YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
MPF-AD + LUT-AD
This repository hosts code for two related projects on weld defect detection:
- MPF-AD (this repo's main contribution) β Parameter-Free Multi-Photometric Fusion for Zero-Shot 3D Weld Anomaly Detection. Manuscript in preparation.
- LUT-AD (parent project, code-only) β Physically Inspired Geometry-to-Photometry Rendering for Real-Time Weld Defect Detection. Accepted at EAAI, 2026.
The two share the same source weld scans and rendering pipeline; the
README below describes the LUT-AD detector code. For the MPF-AD
zero-shot pipeline, see src/pointad_plus/,
scripts/pointad_plus/,
scripts/mvtec3d/,
tests/pointad_plus/,
docs/superpowers/, and the
MPW-AD dataset card.
To reproduce the MPF-AD results (P1βP4 on MPW-AD), follow
docs/reproduce_mpf_ad.md. It covers the
environment (requirements-mpfad.txt), the upstream PointAD commit and
checkpoint, the MPW-AD download, data conversion, the three inference runs
and the analysis scripts. Verified from a fresh clone on a single 8 GB GPU
(about 1.7 h of inference in total).
results/, external/, and Dataset_3D/ are intentionally excluded
from this repo (large data, upstream code, regenerable artefacts).
Geometry-to-Photometry Rendering for Real-Time Weld Defect Detection
Code and dataset for the LUT-AD paper "Physically Inspired Geometry-to-Photometry Rendering for Real-Time Weld Defect Detection" (accepted at Engineering Applications of Artificial Intelligence, EAAI, 2026).
[ δΈζη | Algorithm details | Dataset card | 3D-AD variants ]
Overview
We target weld defect inspection on the top covers of lithium-ion battery cells. Pure 2D imaging captures rich texture but is unreliable on specular metallic surfaces; 3D point-cloud methods capture geometry faithfully but are too heavy for on-line deployment. This work bridges the two by rendering 3D depth into a discriminative 2D photometric image and feeding that to a lightweight 2D detector:
- Depth-Normal Guided Illumination Rendering (C++, OpenCV, Eigen)
- Statistical-prior LUT mapping turns the raw depth map into a pseudo-color image that preserves the global geometric hierarchy.
- Surface normals computed from the depth map drive a Phong shading pass that amplifies subtle local geometric perturbations into high-contrast photometric cues.
- YOLO-WT (YOLO-WaveletTiny) (Python, PyTorch, Ultralytics fork)
- WDSConv in the backbone: depthwise-separable convolution whose 3Γ3 depthwise core is replaced by a multi-level Haar wavelet transform.
- IWUpsample in the neck: replaces interpolation upsampling with multi-level inverse wavelet reconstruction so high-frequency defect cues survive the resolution recovery.
On a real industrial dataset of six battery weld defect classes, the rendered modality alone improves the YOLOv11n baseline by +4.9 mAP@50, and the full YOLO-WT achieves 93.8 % mAP@50 with only 2.47 M parameters and 5.7 GFLOPs.
See docs/algorithm.md for the full mathematical description.
Repository layout
.
βββ Crop_Data/ # Raw cropped 32-bit float depth maps + YOLO labels
β βββ Tif/ # *.tif, single-channel CV_32FC1
β βββ label/ # *.txt, YOLO format (class cx cy w h, normalized)
β
βββ Depth-Normal_Rendering/ # Stage 1: C++ renderer (OpenCV + Eigen)
β βββ LUT.{h,cpp} # Statistical-prior depth β pseudo-color LUT
β βββ Phong.{h,cpp} # Surface normals + Phong shading
β βββ Demo.cpp # Reference pipeline driver
β
βββ Train_Data/ # Stage 1 outputs, organized for detector training
β βββ DepthMap/ # Min-max normalized depth (8-bit)
β βββ LUT/ # Pseudo-color image only (Stage 1a)
β βββ Diffuse/ # Phong diffuse component only
β βββ Specular/ # Phong specular component only
β βββ Normal/ # Surface normal map (color-coded)
β βββ Phong/ # Final rendered image (Stage 1a Γ Stage 1b)
β
βββ YOLO-WT/ # Stage 2: detector
β βββ train.py / val.py / detect.py
β βββ ultralytics/ # Forked Ultralytics tree
β βββ nn/AddModules/ # WDSConv.py, IWUpSample.py
β βββ cfg/models/YOLO-WT/ # YOLO-WT.yaml + WDSConv/IWUpSample ablations
β βββ cfg/datasets/ # One YAML per modality (Phong / LUT / Normal / ...)
β βββ Abl_Exp/train/ # args.yaml + results.csv of the paper's ablation runs
β βββ ANS/ DataPrepare/ Data_Vis/ # Analysis, data-prep and plotting scripts
β
βββ docs/ # Algorithm details and dataset card
The dataset is also released on Hugging Face β see docs/dataset_card.md.
Quick start
1. Clone and fetch the dataset
git clone https://github.com/WillPanSUTD/MPF-AD.git
cd MPF-AD
# Option A: download the dataset from Hugging Face
hf download vpan1226/MPW-AD --repo-type dataset --local-dir .
# Option B: re-render Train_Data from Crop_Data yourself (see Stage 1 below)
1b. Download pre-trained weights
Run from the repository root; files land in ./checkpoints/, which is
where YOLO-WT/val.py and YOLO-WT/detect.py look for them.
# Public release checkpoint (seed 42, val mAP@50 = 0.929)
hf download vpan1226/MPF-AD checkpoints/YOLO-WT-seed42-best.pt --local-dir .
# Original paper run (val mAP@50 = 0.938, the number reported in the paper)
hf download vpan1226/MPF-AD checkpoints/YOLO-WT-paper-best.pt --local-dir .
# Everything: paper run, 4 seed runs, and the WDSConv-only / IWUpSample-only ablations
hf download vpan1226/MPF-AD --include "checkpoints/*" --local-dir .
| File | Run | val mAP@50 |
|---|---|---|
checkpoints/YOLO-WT-paper-best.pt |
Original paper run (seed 0) | 0.938 |
checkpoints/YOLO-WT-seed42-best.pt |
Re-run, seed 42 (public default) | 0.929 |
checkpoints/YOLO-WT-seed{0,1,2}-best.pt |
Re-runs, seeds 0 / 1 / 2 | 0.881 / 0.893 / 0.905 |
checkpoints/ablation/WDSConv-only-best.pt |
Ablation: WDSConv only | 0.921 |
checkpoints/ablation/IWUpSample-only-best.pt |
Ablation: IWUpSample only | 0.932 |
All checkpoints: https://huggingface.co/vpan1226/MPF-AD/tree/main/checkpoints
2. Stage 1 β render depth maps to photometric images (C++)
Dependencies
| Library | Version tested |
|---|---|
| OpenCV | 4.11.0 |
| Eigen | 3.4.0 |
| Compiler | MSVC 2022 / GCC 11+ (C++17) |
Windows (MSVC, included .sln)
- Install OpenCV 4.x and Eigen 3 and add them to the project's include / library paths.
- Open
Depth-Normal_Rendering/Depth-Normal_Rendering.slnin Visual Studio 2022. - Build in
Release / x64and run. Edit the input/output paths at the top ofDemo.cppto point at yourCrop_Data/Tif/*.tif.
Linux (CMake, suggested)
cd Depth-Normal_Rendering
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j
./build/depth_normal_render <depth.tif> <out_dir>
The driver emits, per input depth map: lut.bmp (Stage 1a pseudo-color),
normals.bmp (visualized normals) and phong.bmp (final rendered image).
Default rendering parameters used in the paper:
| Param | Value | Meaning |
|---|---|---|
L, V |
[500, 150, 1500] |
Light / view direction (image space) |
I_l |
1.3 |
Overall light intensity |
w_a |
0.1 |
Ambient weight |
w_d |
0.5 |
Diffuse weight |
w_s |
0.3 |
Specular weight |
n (shine) |
32 |
Specular exponent |
3. Stage 2 β train, validate, predict (Python)
Dependencies
conda create -n yolo-wt python=3.10
conda activate yolo-wt
pip install torch==2.5.1 torchvision --index-url https://download.pytorch.org/whl/cu118
cd YOLO-WT
pip install -r requirements.txt
requirements.txt installs the upstream ultralytics wheel (for its
dependency tree β numpy, opencv-python, tqdm, β¦) plus PyWavelets
and polars. At runtime the bundled fork under YOLO-WT/ultralytics/
takes precedence because the launch directory is at the front of
sys.path, so always cd YOLO-WT before running train.py / val.py.
Train YOLO-WT on the rendered modality
Open YOLO-WT/ultralytics/cfg/datasets/Phong.yaml and set path: / train: /
val: / test: to absolute paths under your local Train_Data/Phong/.
Absolute paths are required here β Ultralytics' dataset loader does not always
resolve relative paths reliably across working directories, so this repo
ships absolute paths as the canonical form. Then:
cd YOLO-WT
python train.py
train.py is configured for the main paper run: model
cfg/models/YOLO-WT/YOLO-WT.yaml, data cfg/datasets/Phong.yaml, 250 epochs,
batch 16, image size 640, SGD, no pretrained weights.
Validate
python val.py # loads ../checkpoints/YOLO-WT-seed42-best.pt
Predict on a single image
python detect.py # edit `source=` to your image
4. Reproducing paper Table results
The dataset YAMLs under cfg/datasets/ (DepthMap.yaml, LUT.yaml, Diffuse.yaml,
Specular.yaml, Normal.yaml, Phong.yaml) point to the six modalities. Re-running
train.py with each YAML reproduces the modality-comparison ablation in the paper
(Sec. 4 / Table β see manuscript). The WDSConv.yaml and IWUpSample.yaml model
configs reproduce the per-module ablation rows.
Hardware used in the paper
| Component | Spec |
|---|---|
| GPU | NVIDIA GeForce RTX 3080 |
| CPU | Intel Core i5-13400F |
| Sensor | OPT-LPC61-3D structured-light line scanner, mounted at 30Β°β45Β° tilt |
Reproducibility
The headline 93.8 % mAP@50 reported in the paper falls comfortably inside the
natural seed variance of training on this small dataset. We retrained the model
under four random seeds with the exact same configuration as the paper
(YOLO-WT.yaml, 250 epochs, batch 16, image size 640, SGD, no pretrained
weights) and validated the saved best.pt of each run with ultralytics' own
validator:
| seed | val P | val R | val mAP@50 | val mAP@50-95 | test mAP@50 |
|---|---|---|---|---|---|
| 0 | 0.859 | 0.822 | 0.881 | 0.467 | 0.773 |
| 1 | 0.877 | 0.794 | 0.893 | 0.471 | 0.749 |
| 2 | 0.831 | 0.866 | 0.905 | 0.469 | 0.774 |
| 42 | 0.847 | 0.879 | 0.929 | 0.480 | 0.788 |
| Paper | β | β | 0.938 | β | β |
Mean of 4 seeds val mAP@50 = 0.902, range = 4.8 pt. Seed 42 is the closest
single-seed reproduction of the paper number and is shipped as the public
release checkpoint on Hugging Face at
vpan1226/MPF-AD under
checkpoints/YOLO-WT-seed42-best.pt. The other three seed runs are kept
for reference and to make the variance verifiable. The original paper run
itself is also released as checkpoints/YOLO-WT-paper-best.pt.
Robustness check via 4-seed Weighted Box Fusion
As a sanity check, we ran a 4-seed Weighted Box Fusion ensemble
(ensemble_eval.py, not yet included in this release; dependencies: pip install ensemble-boxes torchmetrics faster-coco-eval). All metrics here are computed via
torchmetrics.MeanAveragePrecision(backend='faster_coco_eval'), so they are
only directly comparable within this section (they differ slightly from
ultralytics' implementation in the table above):
| val mAP@50 | test mAP@50 | |
|---|---|---|
| seed 42 alone | 0.894 | 0.783 |
| 4-seed WBF ensemble | 0.931 | 0.809 |
The ensemble lifts every class except Bump (which gets pulled down by the weaker seeds' low-confidence boxes); we still ship the single seed-42 checkpoint for ease of deployment. The ensemble result (script to be added in a follow-up release) confirms the single-seed numbers are stable rather than degenerate.
Known limitations of the released checkpoint
- Non-stratified random split. The paper's 8:1:1 split is at the image
level, not stratified by class. Per-class instance counts:
Pseudo
170 / 21 / 23, Pinhole306 / 40 / 35, Pit199 / 22 / 30, Burst76 / 21 / 13, Fish-scale151 / 13 / 24, Bump84 / 6 / 14(train / val / test). With only 6 Bump val instances, single prediction errors swing val Bump AP by 15+ pt; the val number for that class should not be over-interpreted. - Bump β Burst confusion. Both classes are defined as "protrusion β₯ 0.2 mm" with the only distinction being qualitative shape (volumetric vs. fish-scale-like). On the test split, 4 of 14 Bump instances are misclassified as Burst, dragging test Bump mAP to 0.50 (vs. 0.88 on val for the same checkpoint). This appears to be inherent class similarity rather than a training defect; merging the two into a single Protrusion category, or hard-mining BumpβBurst pairs in training, are reasonable future directions.
- Val β test gap of ~12 pt holds across all four seeds and is dominated by the Bump issue above. Treat val numbers as an upper bound of the current dataset's generalization signal, not as deployment performance.
Citation
@article{cao2026geo2pho,
title = {Physically Inspired Geometry-to-Photometry Rendering for Real-Time Weld Defect Detection},
author = {Cao, Ling and Qiu, Jiajun and Zhang, Yunzhi and Feng, Daquan and Pan, Wei},
journal = {Engineering Applications of Artificial Intelligence},
year = {2026},
note = {Accepted, in press. Preprint: SSRN, doi:10.2139/ssrn.6946138}
}
(Volume / article number / DOI will be added once the article is published online.)
Acknowledgments
- Detector built on top of Ultralytics YOLO (AGPL-3.0).
- WDSConv adapts the wavelet convolution idea from Finder et al., WTConv (2024).
- Funded by OPT Machine Vision under the Dongguan Key R&D Program (No. 20241200300122) and the Guangdong Provincial Key R&D Program (No. 2025B0101120001).
License
- Code under
YOLO-WT/ultralytics/: AGPL-3.0, inherited from upstream Ultralytics. - Code under
Depth-Normal_Rendering/: Apache-2.0 (this work). - Dataset under
Crop_Data/andTrain_Data/: seedocs/dataset_card.md.