WaveSeg
WaveSeg is a SegFormer-B0 backbone with a lightweight (770-parameter) Frequency-Boundary Adapter (FBA) gating its decoder features via a Haar wavelet high-frequency branch. It is evaluated on two 2D medical image segmentation domains: brain MRI lesion segmentation and polyp segmentation.
Locked configuration: WaveSeg = SegFormer-B0 + FBA frequency gating (ablation config A1), with NO boundary-frequency auxiliary loss. A boundary-frequency loss was also designed and tested at several weights - it is reported as a negative ablation (see below), not part of the shipped model.
Task and datasets
- Task: binary semantic segmentation, 256x256 input.
- Brain MRI (LGG): FLAIR-abnormality segmentation on the
mateuszbuda/lgg-mri-segmentationdataset (Buda et al., Computers in Biology and Medicine, 2019) - 110 patients, 3,929 slices, patient-level 70/15/15 split. - Polyp (Kvasir-SEG): colonoscopy polyp segmentation on Kvasir-SEG (Jha et al., MMM 2020) - 1,000 images, 800/100/100 split.
Results
All numbers below are read directly from results/ablation_table_clean.csv (grouped/averaged by
method+dataset across all completed runs - n is how many). Region metrics (Dice, mIoU,
Boundary-IoU) are higher-is-better; HD95/ASSD (pixel distances at 256x256) are lower-is-better.
Brain MRI (LGG)
| Method | n | Dice | mIoU | Boundary-IoU | HD95 (px) | ASSD (px) | Params |
|---|---|---|---|---|---|---|---|
| SegFormer-B0 baseline (A0) | 3 | 0.8677 +/- 0.0103 | 0.9159 | 0.7898 | 29.70 | 27.07 | 3,714,401 |
| WaveSeg (A1, ours) | 2 | 0.8908 +/- 0.0055 | 0.9273 | 0.8112 | 22.83 | 20.56 | 3,715,171 |
Kvasir-SEG (Polyp)
| Method | n | Dice | mIoU | Boundary-IoU | HD95 (px) | ASSD (px) | Params |
|---|---|---|---|---|---|---|---|
| SegFormer-B0 baseline (A0) | 1 | 0.9049 | 0.9038 | 0.5809 | 18.49 | 4.57 | 3,714,401 |
| WaveSeg (A1, ours) | 2 | 0.9103 +/- 0.0049 | 0.9084 | 0.5958 | 17.41 | 4.34 | 3,715,171 |
WaveSeg (A1) beats the SegFormer-B0 baseline on every metric, on both datasets, adding only 770 trainable parameters (3,715,171 vs 3,714,401).
Honest negative ablation: the boundary-frequency loss does not help
The FBA module was originally paired with an auxiliary boundary-frequency loss (BCE+Dice
supervising the attention map toward the ground-truth mask's morphological-gradient edge). Tested
on LGG at lambda_boundary in {0.1, 0.25, 0.5, 1.0}, after seed-averaging (n>=2 where re-run):
| lambda_boundary | Dice (LGG) | Beats A1 (gating alone)? |
|---|---|---|
| 0.0 (= A1, shipped) | 0.8908 +/- 0.0055 | reference |
| 0.1 | 0.8826 +/- 0.0174 | no |
| 0.25 | 0.8727 | no |
| 0.5 | 0.8783 | no |
| 1.0 | 0.8753 +/- 0.0069 | no |
No tested weight improved on FBA gating alone. This is reported as a deliberate, honest negative
ablation rather than dropped: the frequency-domain gating mechanism is doing the work; explicitly
supervising its attention map toward GT edges does not add value at any weight tested. See the
training repository's results/tables.md and REPORT.md for the full ablation grid (A0-A5) and
the seed-confirmation methodology.
Training details
- Backbone:
nvidia/mit-b0(ImageNet-pretrained SegFormer-B0 encoder), fresh decode head. - Hardware: local development/inference on an RTX 4060 (8GB); full training runs on Kaggle T4/P100 and locally, all under an 8GB VRAM budget.
- Precision: AMP (mixed precision), channels-last memory format, batch size 16, 256x256 input.
- Optimizer: AdamW, lr 6e-5, weight decay 0.01, 100 epochs.
- Reproducibility: base seed 42 everywhere (data splits fixed across all seed variants);
multi-seed confirmation runs (seeds 42/123/456 depending on the row) give
n>=2for every number reported above except the two Kvasir/LGG A0 rows still atn=1. - Metrics: Dice, mIoU (mean of foreground+background IoU), Boundary-IoU (Cheng et al., CVPR
2021), HD95 and ASSD (via
scipy.ndimage; a fixed image-diagonal penalty is used for the empty-mask edge case so every method aggregates over the identical image set).
Intended use
Research and educational demonstration of frequency-domain gating for medical image segmentation boundary quality. Not validated for clinical use - trained on two small public research datasets (3,929 and 1,000 images respectively), with no independent multi-site validation.
Limitations
- Both datasets are 2D, single-institution(-family) research collections; generalization to other scanners, contrast protocols, or endoscopy hardware is untested.
- The
n=1rows above (Kvasir/LGG A0 baseline on Kvasir) have no seed-variance estimate. - The boundary-frequency loss's negative result was tested only on LGG and only up to
lambda_boundary=1.0; it does not rule out a different loss formulation working. - Small lesions/polyps (a few hundred pixels) are the most common failure mode for both the baseline and WaveSeg.
How to use
See the project README for full reproduction instructions (training, evaluation, and the Gradio demo). The demo Space loads these same weights: https://huggingface.co/spaces/Sarvarbek13/WaveSeg-demo.
Citation
If you use this model, please cite the datasets and metrics it builds on:
@inproceedings{xie2021segformer,
title={SegFormer: Simple and efficient design for semantic segmentation with transformers},
author={Xie, Enze and Wang, Wenhai and Yu, Zhiding and Anandkumar, Anima and Alvarez, Jose M and Luo, Ping},
booktitle={Advances in Neural Information Processing Systems},
volume={34},
pages={12077--12090},
year={2021}
}
@article{buda2019association,
title={Association of genomic subtypes of lower-grade gliomas with shape features automatically extracted by a deep learning algorithm},
author={Buda, Mateusz and Saha, Ashirbani and Mazurowski, Maciej A},
journal={Computers in Biology and Medicine},
volume={109},
pages={218--225},
year={2019},
publisher={Elsevier}
}
@inproceedings{jha2020kvasir,
title={Kvasir-SEG: A segmented polyp dataset},
author={Jha, Debesh and Smedsrud, Pia H and Riegler, Michael A and Halvorsen, P{\aa}l and Lange, Thomas de and Johansen, Dag and Johansen, H{\aa}vard D},
booktitle={International Conference on Multimedia Modeling},
pages={451--462},
year={2020},
organization={Springer}
}
@inproceedings{cheng2021boundary,
title={Boundary IoU: Improving object-centric image segmentation evaluation},
author={Cheng, Bowen and Girshick, Ross and Doll{\'a}r, Piotr and Berg, Alexander C and Kirillov, Alexander},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
pages={15334--15342},
year={2021}
}
Author
Sarvarbek Erniyazov. Licensed under the MIT License (see LICENSE).