YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
- PhD Research Progress Summary
- Medical Image Segmentation on Intel Gaudi HPUs with nnUNet
- Session 1: Data Preparation & Repository Setup
- Session 2: Architecture Changes (50-Epoch Benchmarks)
- Session 3: Inference Optimization
- Cross-Framework Comparison (ACDC β Dataset927)
- Overall Best Results Achieved
- Infrastructure Notes
- Session 4: Ablation Study β Feature Combinations (March 14, 2026)
- Session 5: 1000-Epoch Full Training & Combo Ablation Results (May 3, 2026)
- Overview
- Experiment Scale
- Top 20 Experiments by Validation Dice (50-epoch, fold 0)
- Baseline Full Training Results (1000 epochs, 5-fold CV, XL-plan)
- Completed 1000-Epoch Experiments (Test Inference)
- 1000-Epoch Training Status (12 XL-plan experiments β STOPPED, now resuming)
- Training Times
- Optimizer Configuration Investigation
- What Worked (Key Insights)
- What Didn't Work
- Generalization Status
- Medical Image Segmentation on Intel Gaudi HPUs with nnUNet
PhD Research Progress Summary
Medical Image Segmentation on Intel Gaudi HPUs with nnUNet
Last Updated: 2026-05-03
Dataset: ACDC (Dataset927_ACDC) β Cardiac MRI segmentation
Classes: RV (Right Ventricle), MLV (Myocardium), LVC (Left Ventricle Cavity)
Base Model: nnUNet v2 ResEncUNetXL (3D full resolution, fold 0)
Session 1: Data Preparation & Repository Setup
What Was Done
- Established three benchmark datasets: KiTS2023 (kidney), AMOS2022 (multi-organ), ACDC (cardiac)
- Set up dual framework comparison: nnUNet v2 vs Advanced_nnUNet
- Created result analysis tooling (
summarize_nnunet_results.pyβ 750+ line Excel pivot table generator)
Critical Issues Found & Fixed
| Issue | Problem | Impact | Resolution |
|---|---|---|---|
| Dataset.json format | Wrong label format ("kidney": [1,2,3] vs "0": "background") |
Wrong channel count, incorrect evaluation | β Fixed all 3 datasets |
| Dataset927 metadata | Advanced_nnUNet had kidney labels on cardiac data | Invalidated early comparisons | β Fixed metadata; data files verified identical (MD5) |
| Dataset size mismatch | KiTS: 391 vs 489 samples; AMOS: 288 vs 360, different class counts | Cross-framework comparison caveats | β οΈ Noted β ACDC is identical across both |
Key Conclusions
- Dataset927 (ACDC) is the only dataset with truly identical data + splits across both frameworks β only valid apples-to-apples comparison
- Dataset920 (KiTS2023) has same subset but nnUnet has 25% more data β comparison valid with caveat
- Dataset923 (AMOS) has different class counts (13 vs 15) β not directly comparable
Session 2: Architecture Changes (50-Epoch Benchmarks)
Experimental Setup
- Protocol: 50-epoch screening runs on Dataset927_ACDC, fold 0, 3D full resolution
- Plan: nnUNetResEncUNetXLPlans
- Goal: Identify best architectural modifications before full 1000-epoch training
Completed Results (Ranked by Mean Dice)
| Rank | Trainer | Mean Dice | RV | MLV | LVC | Ξ vs Baseline |
|---|---|---|---|---|---|---|
| 1 | LabelSmooth | 0.9294 | 0.9339 | 0.9047 | 0.9495 | +1.00% |
| 2 | FullKernel | 0.9280 | 0.9343 | 0.9007 | 0.9491 | +0.86% |
| 3 | DiceHeavy | 0.9275 | 0.9300 | 0.9036 | 0.9489 | +0.81% |
| 4 | FatDecoder | 0.9273 | 0.9339 | 0.9057 | 0.9424 | +0.79% |
| 5 | ResDenseUNet | 0.9268 | 0.9321 | 0.9033 | 0.9450 | +0.74% |
| 6 | AdamW_DiceHeavy | 0.9262 | 0.9320 | 0.9040 | 0.9427 | +0.68% |
| 7 | AdamW_TopK | 0.9260 | 0.9335 | 0.9023 | 0.9423 | +0.66% |
| 8 | AdamW3e4 | 0.9247 | 0.9295 | 0.8973 | 0.9473 | +0.53% |
| 9 | DiceTopK10 | 0.9248 | 0.9246 | 0.9012 | 0.9486 | +0.54% |
| 10 | BdryWeight | completed | β | β | β | needs eval |
| 11 | TrueDenseResidual | 0.9197 | 0.9245 | 0.8884 | 0.9464 | +0.03% |
| 12 | ResEncXL (Baseline) | 0.9194 | 0.9238 | 0.8904 | 0.9438 | β |
Still Running (as of 2026-03-14 morning)
| Trainer | Last Epoch | EMA Pseudo Dice | Status |
|---|---|---|---|
| AdamW | 8/50 | 0.8738 | π Running |
| CosWR | 9/50 | 0.5309 | π Running |
Crashed at Epoch 0 (need bug fixes)
| Trainer | Root Cause |
|---|---|
| DropPath | stochastic_depth_p not recognized by dynamic_network_architectures |
| SE_DropPath | Same as DropPath |
| SE_DropPath_DeepDec | Same as DropPath + decoder depth mismatch |
| SE_DropPath_DeepDec_ClassWt | Same as above |
| DeeperDec | Decoder depth parameter issue |
| Combined | Circular import + decoder shape mismatch |
| AttentionUNet | Custom architecture instantiation failure, pydoc.locate() import |
Failed Plan Variants (data incompatibility with ACDC)
| Plan | Mean Dice | Issue |
|---|---|---|
| FineSpacing | 0.1571 | Preprocessing mismatch |
| FineSpacingNNLabel | 0.1511 | Preprocessing mismatch |
| LowOrderData | 0.1858 | Preprocessing mismatch |
| NNLabel | 0.1124 | Preprocessing mismatch |
| Default 50ep baseline | 0.0980 | Training issue (unknown) |
| FocalTversky | 0.0000 | All-zero predictions |
Key Observations
- LabelSmooth is the clear winner (+1.0% over baseline) β simple regularization very effective
- FullKernel (larger convolution kernels) and DiceHeavy loss close behind
- FatDecoder (wider decoder) and ResDenseUNet (hybrid architecture) show architecture changes work
- AdamW optimizer consistently helps across loss variants
- All successful trainers cluster in 91.9%β92.9% β ceiling effect at ~93%
- 7 architectural experiments crashed and need code fixes before re-running
- Plan variants fail catastrophically on ACDC β unnecessary for this dataset
- Performance ceiling estimated at ~0.93 Dice for fold 0, 50-epoch screening
Session 3: Inference Optimization
Baseline
Model: nnUNetTrainer__nnUNetResEncUNetXLPlans__3d_fullres, fold 0
Default inference Dice: 0.9111 (checkpoint_final)
Round 1: Basic Inference Modifications (7 experiments)
| Experiment | Change | Dice | Verdict |
|---|---|---|---|
| E1: checkpoint_best | Use best instead of final checkpoint | 0.9126 (+0.15%) | β KEEP |
| E2: 75% overlap | Sliding window step 0.25 | 0.9111 (Β±0.00%) | β No effect |
| E4: Connected components | Remove small fragments | 0.9111 (Β±0.00%) | β No effect |
| E5: Rotation TTA | 90Β°/180Β°/270Β° rotations | 0.9110 (-0.01%) | β Slight harm |
| E6: Threshold optimization | Per-class probability thresholds | 0.9115 (+0.04%) | β Marginal |
| E7: Cascade 2-stage | Organβsubstructure segmentation | 0.5462 (-36.5%) | β CATASTROPHIC |
Round 2: Advanced Single-Model Techniques (8 experiments)
Baseline: checkpoint_best = 0.9126
| Experiment | Dice | Verdict |
|---|---|---|
| R2E1: Temperature scaling | 0.9126 (Β±0.00%) | β No effect |
| R2E2: Intensity TTA | 0.9126 (Β±0.00%) | β No effect |
| R2E3: Multi-scale TTA | FAILED (bug) | β |
| R2E4: Gaussian sigma tuning | 0.9126 (Β±0.00%) | β No effect |
| R2E5: Probability smoothing | 0.9116 (-0.10%) | β Degradation |
| R2E6: Morphological PP | 0.9116 (-0.10%) | β Degradation |
| R2E7: Anatomical constraints | 0.9139 (+0.13%) | β KEEP |
| R2E8: Multi-checkpoint ensemble | 0.9076 (-0.50%) | β Hurts |
Round 3: Multi-Fold Ensemble (2 experiments)
| Experiment | Dice | RV | MLV | LVC |
|---|---|---|---|---|
| R3E1: 5-fold ensemble | 0.9151 | 0.9058 | 0.8979 | 0.9417 |
| R3E2: 5-fold + anatomy | 0.9161 | 0.9085 | 0.8977 | 0.9421 |
Round 4: Targeted Post-Processing (12 experiments)
Baseline: fold 0 best + anatomy = 0.9139
| Best | Experiment | Dice | Ξ |
|---|---|---|---|
| π | R4E12: CC100 + Bnd092 Anat | 0.9142 | +0.03% |
| β | All other variants | β€0.9139 | Β±0.01% |
| β | R4E10: Smooth + Bnd Slice | 0.8265 | -9.74% |
Round 4b: Step Size Optimization
Result: All step sizes (25%, default, 125%, 375%) produce identical Dice (0.9143) β ACDC volumes too small for step size to matter.
Best Inference Configurations
| Config | Dice | Use Case |
|---|---|---|
| 5-fold ensemble + anatomy | 0.9161 | π Best overall (requires all 5 folds) |
| Fold 0 checkpoint_best | 0.9143 | Fastest single-model |
| Fold 0 best + anatomy | 0.9139 | Single-model with PP |
| Default inference | 0.9111 | Baseline |
Key Inference Lessons
- nnUNet inference is already near-optimal β most modifications yield 0.00% change
- Checkpoint selection matters (+0.15%) β minor late-epoch overfitting exists
- Only domain-specific postprocessing helps β anatomical constraints (+0.13%)
- Multi-fold ensemble is the biggest win (+0.40% over default, +0.18% over tuned single)
- Generic techniques universally fail: temperature scaling, intensity TTA, probability smoothing all ineffective
- Persistent failure cases (patient049, patient034, patient091) are model-inherent β need architecture fixes
- Cascading without training-time support destroys performance (-36.5%)
Cross-Framework Comparison (ACDC β Dataset927)
| Method | RV | MLV | LVC | Mean Dice |
|---|---|---|---|---|
| nnUNet v2 ResEncUNetXL | 0.8902 | 0.9100 | 0.9746 | 0.9249 |
| Advanced ResidualUNet | 0.9058 | 0.9121 | 0.9743 | 0.9307 |
| Advanced CSAttentionUNet | 0.9029 | 0.9121 | 0.9716 | 0.9289 |
Takeaway: Advanced_nnUNet models outperform nnUNet v2 by +0.58% mean Dice, primarily from better RV segmentation. Attention/hybrid architectures handle the irregular RV shape better.
Overall Best Results Achieved
| Metric | Value | Configuration |
|---|---|---|
| Best 50-ep screening | 0.9294 | LabelSmooth trainer |
| Best inference (single model) | 0.9143 | checkpoint_best, default step |
| Best inference (ensemble) | 0.9161 | 5-fold ensemble + anatomy |
| Best cross-framework | 0.9307 | Advanced_nnUNet ResidualUNet |
Infrastructure Notes
Pod Management
- Always create a new pod for each new experiment β reusing pods across different experiments leads to evictions and wasted GPU-hours
- Same pod OK for: continuing/resuming the same experiment, re-running validation on same model
- New pod needed for: different trainer, different dataset, different architecture
- Pod template:
/home/kkaczor/Desktop/scripts/moje/yamls/124_latest_2100_phd.yaml - Create:
hlctl create containers --flavor g3 --file <yaml> --name kkaczor-<task> --namespace framework --username kkaczor -q --no-pager --as-job --shm 10240 - Fresh pods need:
pip install -e /software/users/kkaczor/phd/nnUnet/nnUNet - Data: shared weka storage at
/software/users/kkaczor/phd/nnUnet/
Environment Variables (set before training)
export nnUNet_raw="/software/users/kkaczor/phd/nnUnet/nnUNet_raw"
export nnUNet_preprocessed="/software/users/kkaczor/phd/nnUnet/preprocessed"
export nnUNet_results="/software/users/kkaczor/phd/nnUnet/results"
Session 4: Ablation Study β Feature Combinations (March 14, 2026)
Design Rationale
From the 50-epoch benchmarks, 4 independent successful features were identified on the ResEncUNetXL architecture, plus ResDenseUNet as a standalone:
| Code | Feature | Solo Dice | Effect |
|---|---|---|---|
| LS | Label Smoothing 0.1 | 0.9294 | Loss regularization β prevents overconfident predictions |
| FK | Full [3,3,3] Kernels | 0.9280 | Architecture β through-plane context at all stages |
| DH | Dice-Heavy (2:0.5) | 0.9275 | Loss β prioritizes Dice metric over CE |
| FD | Fat Decoder + SE + StochDepth | 0.9273 | Architecture β 3x decoder capacity + attention |
| RD | ResDenseUNet | 0.9268 | Alternative architecture (residual enc + dense dec) |
All combinations use consistent base: AdamW 1e-3 + CosineAnnealingLR + TopK(k=10) loss.
For the ablation-in-pairs approach (AdamW variants 6-9), AdamW_DiceHeavy (0.9262) was the best; since DiceHeavy is already a separate feature (DH), the AdamW optimizer is standardized as the base.
Full Experiment Matrix
Main ablation path (for publication):
Baseline (0.9194) β +LS (0.9294) β +LS+FK β +LS+FK+DH β +LS+FK+DH+FD
All 14 new experiments:
| # | Combination | Type | Features |
|---|---|---|---|
| 1 | ComboLS_FK | Pair | LabelSmooth + FullKernel |
| 2 | ComboLS_FD | Pair | LabelSmooth + FatDecoder |
| 3 | ComboLS_DH | Pair | LabelSmooth + DiceHeavy |
| 4 | ComboFK_FD | Pair | FullKernel + FatDecoder |
| 5 | ComboFK_DH | Pair | FullKernel + DiceHeavy |
| 6 | ComboFD_DH | Pair | FatDecoder + DiceHeavy |
| 7 | ComboRD_LS | Pair | ResDense + LabelSmooth |
| 8 | ComboRD_DH | Pair | ResDense + DiceHeavy |
| 9 | ComboLS_FK_FD | Triple | LS + FK + FD |
| 10 | ComboLS_FK_DH | Triple | LS + FK + DH |
| 11 | ComboLS_FD_DH | Triple | LS + FD + DH |
| 12 | ComboFK_FD_DH | Triple | FK + FD + DH |
| 13 | ComboRD_LS_DH | Triple | RD + LS + DH |
| 14 | ComboLS_FK_FD_DH | Quad | ALL FOUR combined |
Trainer Files
All located in: nnunetv2/training/nnUNetTrainer/variants/combinations/
Launch
bash /home/kkaczor/software/phd/nnUnet/launch_ablation_study.sh
Creates 14 pods (one per experiment), each running 50 epochs on fold 0.
Session 5: 1000-Epoch Full Training & Combo Ablation Results (May 3, 2026)
Overview
The 50-epoch screening phase is essentially complete with 89 experiments finished. The focus has shifted to full 1000-epoch training of the most promising combinations, with 14 experiments promoted from screening.
Experiment Scale
| Category | Count | Status |
|---|---|---|
| 50-epoch screening runs | 89 | β All completed |
| 100-epoch extended screens | 2 | β Completed |
| 1000-epoch full training | 14 | 2 completed, 12 stopped (pod recycled) |
| Multi-fold cross-validation | 3 architectures | Baseline done, others partial |
| Total experiment dirs | 120 | β |
Top 20 Experiments by Validation Dice (50-epoch, fold 0)
| Rank | Experiment | Val Dice | Key Features |
|---|---|---|---|
| 1 | ComboLS_LateZ_DH_DeepDec_Xavier | 0.9329 | LS + LateZ + DiceHeavy + DeepDecoder + Xavier init |
| 2 | ComboLS_LateZ_DH_DeepDec_Reinit50 | 0.9323 | LS + LateZ + DiceHeavy + DeepDecoder + Reinit@50 |
| 3 | ComboLS_LateZ_DH_DeepDec_Xavier_DeeperEnc | 0.9321 | Above + deeper encoder |
| 4 | ComboLS_LateZ_DH_DeepDec_Xavier_DropPath | 0.9314 | Above + stochastic depth |
| 5 | ComboLS_FK_DH | 0.9312 | LS + FullKernel + DiceHeavy |
| 6 | ComboLS_LateZ | 0.9311 | LS + LateZ kernel |
| 7 | ComboLS_LateZ_DH_DeepDec | 0.9310 | LS + LateZ + DiceHeavy + DeepDecoder |
| 8 | ComboLS_FK_FD_DH | 0.9310 | LS + FullKernel + FatDecoder + DiceHeavy |
| 9 | ComboLS_FK_DH (100ep) | 0.9309 | Extended screening |
| 10 | ComboLS_LateZ_DH_DeepDec_Reinit100 | 0.9309 | LS + LateZ + DH + DeepDec + Reinit@100 |
| 11 | ComboFK_FD_DH | 0.9309 | FK + FatDecoder + DiceHeavy |
| 12 | DeepDec (L-plan) | 0.9306 | Deep decoder on smaller model |
| 13 | ComboLS_LateZ_DH_DeepDec_Xavier_TopK5 | 0.9305 | Xavier + TopK5 loss |
| 14 | ComboLS_LateZ_DH_GELU | 0.9304 | LS + LateZ + DH + GELU activation |
| 15 | ComboLS_LateZ_Xavier | 0.9302 | LS + LateZ + Xavier init |
| 16 | ComboLS_FK_DH_Xavier | 0.9300 | LS + FK + DH + Xavier init |
| 17 | ComboLS_LateZ_DH (L-plan) | 0.9300 | Smaller model variant |
| 18 | ComboLS_FK_DH_TopK20 | 0.9299 | LS + FK + DH + TopK20 loss |
| 19 | ComboFK_FD | 0.9299 | FK + FatDecoder |
| 20 | BdryWeight (M-plan) | 0.9297 | Boundary weighting |
50-epoch Baseline (ResEncUNetXL): 0.9194
Baseline Full Training Results (1000 epochs, 5-fold CV, XL-plan)
| Fold | Mean Dice | RV | MLV | LVC |
|---|---|---|---|---|
| fold_0 | 0.9111 | 0.9014 | 0.8930 | 0.9388 |
| fold_1 | 0.9116 | 0.9067 | 0.8898 | 0.9382 |
| fold_2 | 0.9076 | 0.8956 | 0.8902 | 0.9370 |
| fold_3 | 0.9090 | 0.8989 | 0.8913 | 0.9366 |
| fold_4 | 0.9110 | 0.9026 | 0.8918 | 0.9386 |
| Average | 0.9101 | 0.9010 | 0.8912 | 0.9378 |
Completed 1000-Epoch Experiments (Test Inference)
| Experiment | Plan | Fold 0 Test Dice | vs Baseline | RV | MLV | LVC |
|---|---|---|---|---|---|---|
| Baseline | XL | 0.9111 | β | 0.9014 | 0.8930 | 0.9388 |
| TrueDenseResidualUNet | XL | 0.9084 | -0.27% | 0.8962 | 0.8923 | 0.9368 |
| DenseUNet | XL | 0.9034 | -0.77% | 0.9007 | 0.8875 | 0.9222 |
| DeepDec_1000ep | L | 0.9043 | -0.68% | 0.8770 | 0.8949 | 0.9411 |
| ComboLS_LateZ_DH_1000ep | L | 0.8793 | -3.18% | 0.8441 | 0.8680 | 0.9257 |
Note: DeepDec and ComboLS_LateZ_DH used L-plan (smaller model), not XL. Comparison with XL baseline is not apples-to-apples. The XL-plan combo variants are the main challengers.
1000-Epoch Training Status (12 XL-plan experiments β STOPPED, now resuming)
| Experiment | 50ep Dice | Last Epoch | Remaining | Est. Time | Status |
|---|---|---|---|---|---|
| ComboLS_LateZ_DH_DeepDec_Xavier | 0.9329 | 604 | 396 | ~90h | π Resuming |
| ComboLS_LateZ_DH_DeepDec_Reinit50 | 0.9323 | 565 | 435 | ~99h | π Resuming |
| ComboLS_LateZ_DH_DeepDec_Xavier_DeeperEnc | 0.9321 | 581 | 419 | ~95h | π Resuming |
| ComboLS_LateZ_DH_DeepDec_Xavier_DropPath | 0.9314 | 599 | 401 | ~91h | π Resuming |
| ComboLS_LateZ | 0.9311 | 716 | 284 | ~21h | π Resuming |
| ComboLS_LateZ_DH_GELU | 0.9304 | 732 | 268 | ~61h | βΈοΈ Queued |
| ComboLS_FK_FD_DH | 0.9310 | 588 | 412 | ~93h | βΈοΈ Queued |
| ComboFK_FD_DH | 0.9309 | 540 | 460 | ~104h | βΈοΈ Queued |
| ComboLS_LateZ_DH_DeepDec | 0.9310 | 522 | 478 | ~109h | βΈοΈ Queued |
| ComboLS_FK_DH_TopK20 | 0.9299 | 416 | 584 | ~133h | βΈοΈ Queued |
| ComboLS_FK_DH | 0.9312 | 208 | 792 | ~180h | βΈοΈ Queued |
| ComboFK_FD | 0.9299 | 362 | 638 | ~145h | βΈοΈ Queued |
Pod kkaczor-1k-resume-tfjob launched May 3, 2026 on Gaudi 3 (g3) to resume the top 5 experiments.
Training Times
| Configuration | Avg Epoch Time | Est. Total (1000ep) |
|---|---|---|
| Baseline (XL) | 246s (4.1 min) | ~68h |
| ComboLS_LateZ_DH_DeepDec_Xavier (XL) | 818s (13.6 min) | ~227h |
| DeepDec (L-plan) | 140s (2.3 min) | ~39h |
Optimizer Configuration Investigation
All 1000-epoch trainers were investigated for a suspected LR schedule bug (T_max targeting 50 instead of 1000):
- Source code: All 14 trainers override
configure_optimizers()withT_max=self.num_epochs(1000) β - debug.json: All show
num_epochs=1000β - Training logs: LR at epoch 50 = 0.00099 (correct for T_max=1000). If buggy, it would be ~1e-6 β
- Optimizer: torch.optim.AdamW (switched from FusedAdamW to fix step-counter serialization bug on checkpoint resume)
- Settings: initial_lr=1e-3, weight_decay=3e-2, CosineAnnealingLR, eta_min=1e-6
Verdict: LR schedule fix was applied BEFORE any 1000-epoch training started. No experiments affected by the bug.
What Worked (Key Insights)
- Label Smoothing (LS) is the single most impactful regularization technique (+1.0% over baseline at 50ep)
- LateZ kernels (late application of z-axis convolutions) combine very well with LS
- Deep Decoder with Xavier initialization is the top combination (0.9329 at 50ep)
- DiceHeavy loss (2:0.5 Dice:CE ratio) consistently helps across all combinations
- Combinations outperform singles β the ablation path Baselineβ+LSβ+LateZβ+DHβ+DeepDecβ+Xavier shows monotonic improvement
- Reinit strategies (reinitializing weights at epoch 50) work nearly as well as Xavier from scratch
- AdamW + CosineAnnealingLR is the best optimizer setup (vs SGD+PolyLR baseline)
- Model size matters at test time β L-plan 1000ep experiments (DeepDec: 0.9043, ComboLS_LateZ_DH: 0.8793) significantly underperform XL baseline (0.9111)
What Didn't Work
- FatDecoder (FD) β ranked lower in combos than expected; overhead doesn't justify gains
- Alternative architectures β DenseUNet (0.9034) and TrueDenseResidual (0.9084) trail baseline
- Plan variants (FineSpacing, LowOrder, NNLabel) β catastrophic failure on ACDC
- Generic inference postprocessing β temperature scaling, TTA, smoothing all ineffective
- FocalTversky loss β produced all-zero predictions
- GELU activation β marginal improvement, not worth the overhead
- TopK5 loss β too aggressive, slightly hurts
- L-plan models for 1000ep β insufficient capacity; the 3.18% gap of ComboLS_LateZ_DH shows smaller models can't match XL even with better training tricks
Generalization Status
Cross-dataset experiments on Dataset920 (KiTS2023) and Dataset923 (AMOS2022) have partial results. Some crashed and need reruns. This is tracked in launch_all_missing.py.