YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

PhD Research Progress Summary

Medical Image Segmentation on Intel Gaudi HPUs with nnUNet

Last Updated: 2026-05-03
Dataset: ACDC (Dataset927_ACDC) β€” Cardiac MRI segmentation
Classes: RV (Right Ventricle), MLV (Myocardium), LVC (Left Ventricle Cavity)
Base Model: nnUNet v2 ResEncUNetXL (3D full resolution, fold 0)


Session 1: Data Preparation & Repository Setup

What Was Done

  • Established three benchmark datasets: KiTS2023 (kidney), AMOS2022 (multi-organ), ACDC (cardiac)
  • Set up dual framework comparison: nnUNet v2 vs Advanced_nnUNet
  • Created result analysis tooling (summarize_nnunet_results.py β€” 750+ line Excel pivot table generator)

Critical Issues Found & Fixed

Issue Problem Impact Resolution
Dataset.json format Wrong label format ("kidney": [1,2,3] vs "0": "background") Wrong channel count, incorrect evaluation βœ… Fixed all 3 datasets
Dataset927 metadata Advanced_nnUNet had kidney labels on cardiac data Invalidated early comparisons βœ… Fixed metadata; data files verified identical (MD5)
Dataset size mismatch KiTS: 391 vs 489 samples; AMOS: 288 vs 360, different class counts Cross-framework comparison caveats ⚠️ Noted β€” ACDC is identical across both

Key Conclusions

  • Dataset927 (ACDC) is the only dataset with truly identical data + splits across both frameworks β†’ only valid apples-to-apples comparison
  • Dataset920 (KiTS2023) has same subset but nnUnet has 25% more data β†’ comparison valid with caveat
  • Dataset923 (AMOS) has different class counts (13 vs 15) β†’ not directly comparable

Session 2: Architecture Changes (50-Epoch Benchmarks)

Experimental Setup

  • Protocol: 50-epoch screening runs on Dataset927_ACDC, fold 0, 3D full resolution
  • Plan: nnUNetResEncUNetXLPlans
  • Goal: Identify best architectural modifications before full 1000-epoch training

Completed Results (Ranked by Mean Dice)

Rank Trainer Mean Dice RV MLV LVC Ξ” vs Baseline
1 LabelSmooth 0.9294 0.9339 0.9047 0.9495 +1.00%
2 FullKernel 0.9280 0.9343 0.9007 0.9491 +0.86%
3 DiceHeavy 0.9275 0.9300 0.9036 0.9489 +0.81%
4 FatDecoder 0.9273 0.9339 0.9057 0.9424 +0.79%
5 ResDenseUNet 0.9268 0.9321 0.9033 0.9450 +0.74%
6 AdamW_DiceHeavy 0.9262 0.9320 0.9040 0.9427 +0.68%
7 AdamW_TopK 0.9260 0.9335 0.9023 0.9423 +0.66%
8 AdamW3e4 0.9247 0.9295 0.8973 0.9473 +0.53%
9 DiceTopK10 0.9248 0.9246 0.9012 0.9486 +0.54%
10 BdryWeight completed β€” β€” β€” needs eval
11 TrueDenseResidual 0.9197 0.9245 0.8884 0.9464 +0.03%
12 ResEncXL (Baseline) 0.9194 0.9238 0.8904 0.9438 β€”

Still Running (as of 2026-03-14 morning)

Trainer Last Epoch EMA Pseudo Dice Status
AdamW 8/50 0.8738 πŸ”„ Running
CosWR 9/50 0.5309 πŸ”„ Running

Crashed at Epoch 0 (need bug fixes)

Trainer Root Cause
DropPath stochastic_depth_p not recognized by dynamic_network_architectures
SE_DropPath Same as DropPath
SE_DropPath_DeepDec Same as DropPath + decoder depth mismatch
SE_DropPath_DeepDec_ClassWt Same as above
DeeperDec Decoder depth parameter issue
Combined Circular import + decoder shape mismatch
AttentionUNet Custom architecture instantiation failure, pydoc.locate() import

Failed Plan Variants (data incompatibility with ACDC)

Plan Mean Dice Issue
FineSpacing 0.1571 Preprocessing mismatch
FineSpacingNNLabel 0.1511 Preprocessing mismatch
LowOrderData 0.1858 Preprocessing mismatch
NNLabel 0.1124 Preprocessing mismatch
Default 50ep baseline 0.0980 Training issue (unknown)
FocalTversky 0.0000 All-zero predictions

Key Observations

  1. LabelSmooth is the clear winner (+1.0% over baseline) β€” simple regularization very effective
  2. FullKernel (larger convolution kernels) and DiceHeavy loss close behind
  3. FatDecoder (wider decoder) and ResDenseUNet (hybrid architecture) show architecture changes work
  4. AdamW optimizer consistently helps across loss variants
  5. All successful trainers cluster in 91.9%–92.9% β€” ceiling effect at ~93%
  6. 7 architectural experiments crashed and need code fixes before re-running
  7. Plan variants fail catastrophically on ACDC β€” unnecessary for this dataset
  8. Performance ceiling estimated at ~0.93 Dice for fold 0, 50-epoch screening

Session 3: Inference Optimization

Baseline

Model: nnUNetTrainer__nnUNetResEncUNetXLPlans__3d_fullres, fold 0
Default inference Dice: 0.9111 (checkpoint_final)

Round 1: Basic Inference Modifications (7 experiments)

Experiment Change Dice Verdict
E1: checkpoint_best Use best instead of final checkpoint 0.9126 (+0.15%) βœ… KEEP
E2: 75% overlap Sliding window step 0.25 0.9111 (±0.00%) ❌ No effect
E4: Connected components Remove small fragments 0.9111 (±0.00%) ❌ No effect
E5: Rotation TTA 90°/180°/270° rotations 0.9110 (-0.01%) ❌ Slight harm
E6: Threshold optimization Per-class probability thresholds 0.9115 (+0.04%) ❌ Marginal
E7: Cascade 2-stage Organβ†’substructure segmentation 0.5462 (-36.5%) ❌ CATASTROPHIC

Round 2: Advanced Single-Model Techniques (8 experiments)

Baseline: checkpoint_best = 0.9126

Experiment Dice Verdict
R2E1: Temperature scaling 0.9126 (±0.00%) ❌ No effect
R2E2: Intensity TTA 0.9126 (±0.00%) ❌ No effect
R2E3: Multi-scale TTA FAILED (bug) ❌
R2E4: Gaussian sigma tuning 0.9126 (±0.00%) ❌ No effect
R2E5: Probability smoothing 0.9116 (-0.10%) ❌ Degradation
R2E6: Morphological PP 0.9116 (-0.10%) ❌ Degradation
R2E7: Anatomical constraints 0.9139 (+0.13%) βœ… KEEP
R2E8: Multi-checkpoint ensemble 0.9076 (-0.50%) ❌ Hurts

Round 3: Multi-Fold Ensemble (2 experiments)

Experiment Dice RV MLV LVC
R3E1: 5-fold ensemble 0.9151 0.9058 0.8979 0.9417
R3E2: 5-fold + anatomy 0.9161 0.9085 0.8977 0.9421

Round 4: Targeted Post-Processing (12 experiments)

Baseline: fold 0 best + anatomy = 0.9139

Best Experiment Dice Ξ”
πŸ† R4E12: CC100 + Bnd092 Anat 0.9142 +0.03%
β€” All other variants ≀0.9139 Β±0.01%
❌ R4E10: Smooth + Bnd Slice 0.8265 -9.74%

Round 4b: Step Size Optimization

Result: All step sizes (25%, default, 125%, 375%) produce identical Dice (0.9143) β€” ACDC volumes too small for step size to matter.

Best Inference Configurations

Config Dice Use Case
5-fold ensemble + anatomy 0.9161 πŸ† Best overall (requires all 5 folds)
Fold 0 checkpoint_best 0.9143 Fastest single-model
Fold 0 best + anatomy 0.9139 Single-model with PP
Default inference 0.9111 Baseline

Key Inference Lessons

  1. nnUNet inference is already near-optimal β€” most modifications yield 0.00% change
  2. Checkpoint selection matters (+0.15%) β€” minor late-epoch overfitting exists
  3. Only domain-specific postprocessing helps β€” anatomical constraints (+0.13%)
  4. Multi-fold ensemble is the biggest win (+0.40% over default, +0.18% over tuned single)
  5. Generic techniques universally fail: temperature scaling, intensity TTA, probability smoothing all ineffective
  6. Persistent failure cases (patient049, patient034, patient091) are model-inherent β€” need architecture fixes
  7. Cascading without training-time support destroys performance (-36.5%)

Cross-Framework Comparison (ACDC β€” Dataset927)

Method RV MLV LVC Mean Dice
nnUNet v2 ResEncUNetXL 0.8902 0.9100 0.9746 0.9249
Advanced ResidualUNet 0.9058 0.9121 0.9743 0.9307
Advanced CSAttentionUNet 0.9029 0.9121 0.9716 0.9289

Takeaway: Advanced_nnUNet models outperform nnUNet v2 by +0.58% mean Dice, primarily from better RV segmentation. Attention/hybrid architectures handle the irregular RV shape better.


Overall Best Results Achieved

Metric Value Configuration
Best 50-ep screening 0.9294 LabelSmooth trainer
Best inference (single model) 0.9143 checkpoint_best, default step
Best inference (ensemble) 0.9161 5-fold ensemble + anatomy
Best cross-framework 0.9307 Advanced_nnUNet ResidualUNet

Infrastructure Notes

Pod Management

  • Always create a new pod for each new experiment β€” reusing pods across different experiments leads to evictions and wasted GPU-hours
  • Same pod OK for: continuing/resuming the same experiment, re-running validation on same model
  • New pod needed for: different trainer, different dataset, different architecture
  • Pod template: /home/kkaczor/Desktop/scripts/moje/yamls/124_latest_2100_phd.yaml
  • Create: hlctl create containers --flavor g3 --file <yaml> --name kkaczor-<task> --namespace framework --username kkaczor -q --no-pager --as-job --shm 10240
  • Fresh pods need: pip install -e /software/users/kkaczor/phd/nnUnet/nnUNet
  • Data: shared weka storage at /software/users/kkaczor/phd/nnUnet/

Environment Variables (set before training)

export nnUNet_raw="/software/users/kkaczor/phd/nnUnet/nnUNet_raw"
export nnUNet_preprocessed="/software/users/kkaczor/phd/nnUnet/preprocessed"
export nnUNet_results="/software/users/kkaczor/phd/nnUnet/results"

Session 4: Ablation Study β€” Feature Combinations (March 14, 2026)

Design Rationale

From the 50-epoch benchmarks, 4 independent successful features were identified on the ResEncUNetXL architecture, plus ResDenseUNet as a standalone:

Code Feature Solo Dice Effect
LS Label Smoothing 0.1 0.9294 Loss regularization β€” prevents overconfident predictions
FK Full [3,3,3] Kernels 0.9280 Architecture β€” through-plane context at all stages
DH Dice-Heavy (2:0.5) 0.9275 Loss β€” prioritizes Dice metric over CE
FD Fat Decoder + SE + StochDepth 0.9273 Architecture β€” 3x decoder capacity + attention
RD ResDenseUNet 0.9268 Alternative architecture (residual enc + dense dec)

All combinations use consistent base: AdamW 1e-3 + CosineAnnealingLR + TopK(k=10) loss.

For the ablation-in-pairs approach (AdamW variants 6-9), AdamW_DiceHeavy (0.9262) was the best; since DiceHeavy is already a separate feature (DH), the AdamW optimizer is standardized as the base.

Full Experiment Matrix

Main ablation path (for publication):

Baseline (0.9194) β†’ +LS (0.9294) β†’ +LS+FK β†’ +LS+FK+DH β†’ +LS+FK+DH+FD

All 14 new experiments:

# Combination Type Features
1 ComboLS_FK Pair LabelSmooth + FullKernel
2 ComboLS_FD Pair LabelSmooth + FatDecoder
3 ComboLS_DH Pair LabelSmooth + DiceHeavy
4 ComboFK_FD Pair FullKernel + FatDecoder
5 ComboFK_DH Pair FullKernel + DiceHeavy
6 ComboFD_DH Pair FatDecoder + DiceHeavy
7 ComboRD_LS Pair ResDense + LabelSmooth
8 ComboRD_DH Pair ResDense + DiceHeavy
9 ComboLS_FK_FD Triple LS + FK + FD
10 ComboLS_FK_DH Triple LS + FK + DH
11 ComboLS_FD_DH Triple LS + FD + DH
12 ComboFK_FD_DH Triple FK + FD + DH
13 ComboRD_LS_DH Triple RD + LS + DH
14 ComboLS_FK_FD_DH Quad ALL FOUR combined

Trainer Files

All located in: nnunetv2/training/nnUNetTrainer/variants/combinations/

Launch

bash /home/kkaczor/software/phd/nnUnet/launch_ablation_study.sh

Creates 14 pods (one per experiment), each running 50 epochs on fold 0.


Session 5: 1000-Epoch Full Training & Combo Ablation Results (May 3, 2026)

Overview

The 50-epoch screening phase is essentially complete with 89 experiments finished. The focus has shifted to full 1000-epoch training of the most promising combinations, with 14 experiments promoted from screening.

Experiment Scale

Category Count Status
50-epoch screening runs 89 βœ… All completed
100-epoch extended screens 2 βœ… Completed
1000-epoch full training 14 2 completed, 12 stopped (pod recycled)
Multi-fold cross-validation 3 architectures Baseline done, others partial
Total experiment dirs 120 β€”

Top 20 Experiments by Validation Dice (50-epoch, fold 0)

Rank Experiment Val Dice Key Features
1 ComboLS_LateZ_DH_DeepDec_Xavier 0.9329 LS + LateZ + DiceHeavy + DeepDecoder + Xavier init
2 ComboLS_LateZ_DH_DeepDec_Reinit50 0.9323 LS + LateZ + DiceHeavy + DeepDecoder + Reinit@50
3 ComboLS_LateZ_DH_DeepDec_Xavier_DeeperEnc 0.9321 Above + deeper encoder
4 ComboLS_LateZ_DH_DeepDec_Xavier_DropPath 0.9314 Above + stochastic depth
5 ComboLS_FK_DH 0.9312 LS + FullKernel + DiceHeavy
6 ComboLS_LateZ 0.9311 LS + LateZ kernel
7 ComboLS_LateZ_DH_DeepDec 0.9310 LS + LateZ + DiceHeavy + DeepDecoder
8 ComboLS_FK_FD_DH 0.9310 LS + FullKernel + FatDecoder + DiceHeavy
9 ComboLS_FK_DH (100ep) 0.9309 Extended screening
10 ComboLS_LateZ_DH_DeepDec_Reinit100 0.9309 LS + LateZ + DH + DeepDec + Reinit@100
11 ComboFK_FD_DH 0.9309 FK + FatDecoder + DiceHeavy
12 DeepDec (L-plan) 0.9306 Deep decoder on smaller model
13 ComboLS_LateZ_DH_DeepDec_Xavier_TopK5 0.9305 Xavier + TopK5 loss
14 ComboLS_LateZ_DH_GELU 0.9304 LS + LateZ + DH + GELU activation
15 ComboLS_LateZ_Xavier 0.9302 LS + LateZ + Xavier init
16 ComboLS_FK_DH_Xavier 0.9300 LS + FK + DH + Xavier init
17 ComboLS_LateZ_DH (L-plan) 0.9300 Smaller model variant
18 ComboLS_FK_DH_TopK20 0.9299 LS + FK + DH + TopK20 loss
19 ComboFK_FD 0.9299 FK + FatDecoder
20 BdryWeight (M-plan) 0.9297 Boundary weighting

50-epoch Baseline (ResEncUNetXL): 0.9194

Baseline Full Training Results (1000 epochs, 5-fold CV, XL-plan)

Fold Mean Dice RV MLV LVC
fold_0 0.9111 0.9014 0.8930 0.9388
fold_1 0.9116 0.9067 0.8898 0.9382
fold_2 0.9076 0.8956 0.8902 0.9370
fold_3 0.9090 0.8989 0.8913 0.9366
fold_4 0.9110 0.9026 0.8918 0.9386
Average 0.9101 0.9010 0.8912 0.9378

Completed 1000-Epoch Experiments (Test Inference)

Experiment Plan Fold 0 Test Dice vs Baseline RV MLV LVC
Baseline XL 0.9111 β€” 0.9014 0.8930 0.9388
TrueDenseResidualUNet XL 0.9084 -0.27% 0.8962 0.8923 0.9368
DenseUNet XL 0.9034 -0.77% 0.9007 0.8875 0.9222
DeepDec_1000ep L 0.9043 -0.68% 0.8770 0.8949 0.9411
ComboLS_LateZ_DH_1000ep L 0.8793 -3.18% 0.8441 0.8680 0.9257

Note: DeepDec and ComboLS_LateZ_DH used L-plan (smaller model), not XL. Comparison with XL baseline is not apples-to-apples. The XL-plan combo variants are the main challengers.

1000-Epoch Training Status (12 XL-plan experiments β€” STOPPED, now resuming)

Experiment 50ep Dice Last Epoch Remaining Est. Time Status
ComboLS_LateZ_DH_DeepDec_Xavier 0.9329 604 396 ~90h πŸ”„ Resuming
ComboLS_LateZ_DH_DeepDec_Reinit50 0.9323 565 435 ~99h πŸ”„ Resuming
ComboLS_LateZ_DH_DeepDec_Xavier_DeeperEnc 0.9321 581 419 ~95h πŸ”„ Resuming
ComboLS_LateZ_DH_DeepDec_Xavier_DropPath 0.9314 599 401 ~91h πŸ”„ Resuming
ComboLS_LateZ 0.9311 716 284 ~21h πŸ”„ Resuming
ComboLS_LateZ_DH_GELU 0.9304 732 268 ~61h ⏸️ Queued
ComboLS_FK_FD_DH 0.9310 588 412 ~93h ⏸️ Queued
ComboFK_FD_DH 0.9309 540 460 ~104h ⏸️ Queued
ComboLS_LateZ_DH_DeepDec 0.9310 522 478 ~109h ⏸️ Queued
ComboLS_FK_DH_TopK20 0.9299 416 584 ~133h ⏸️ Queued
ComboLS_FK_DH 0.9312 208 792 ~180h ⏸️ Queued
ComboFK_FD 0.9299 362 638 ~145h ⏸️ Queued

Pod kkaczor-1k-resume-tfjob launched May 3, 2026 on Gaudi 3 (g3) to resume the top 5 experiments.

Training Times

Configuration Avg Epoch Time Est. Total (1000ep)
Baseline (XL) 246s (4.1 min) ~68h
ComboLS_LateZ_DH_DeepDec_Xavier (XL) 818s (13.6 min) ~227h
DeepDec (L-plan) 140s (2.3 min) ~39h

Optimizer Configuration Investigation

All 1000-epoch trainers were investigated for a suspected LR schedule bug (T_max targeting 50 instead of 1000):

  • Source code: All 14 trainers override configure_optimizers() with T_max=self.num_epochs (1000) βœ…
  • debug.json: All show num_epochs=1000 βœ…
  • Training logs: LR at epoch 50 = 0.00099 (correct for T_max=1000). If buggy, it would be ~1e-6 βœ…
  • Optimizer: torch.optim.AdamW (switched from FusedAdamW to fix step-counter serialization bug on checkpoint resume)
  • Settings: initial_lr=1e-3, weight_decay=3e-2, CosineAnnealingLR, eta_min=1e-6

Verdict: LR schedule fix was applied BEFORE any 1000-epoch training started. No experiments affected by the bug.

What Worked (Key Insights)

  1. Label Smoothing (LS) is the single most impactful regularization technique (+1.0% over baseline at 50ep)
  2. LateZ kernels (late application of z-axis convolutions) combine very well with LS
  3. Deep Decoder with Xavier initialization is the top combination (0.9329 at 50ep)
  4. DiceHeavy loss (2:0.5 Dice:CE ratio) consistently helps across all combinations
  5. Combinations outperform singles — the ablation path Baseline→+LS→+LateZ→+DH→+DeepDec→+Xavier shows monotonic improvement
  6. Reinit strategies (reinitializing weights at epoch 50) work nearly as well as Xavier from scratch
  7. AdamW + CosineAnnealingLR is the best optimizer setup (vs SGD+PolyLR baseline)
  8. Model size matters at test time β€” L-plan 1000ep experiments (DeepDec: 0.9043, ComboLS_LateZ_DH: 0.8793) significantly underperform XL baseline (0.9111)

What Didn't Work

  1. FatDecoder (FD) β€” ranked lower in combos than expected; overhead doesn't justify gains
  2. Alternative architectures β€” DenseUNet (0.9034) and TrueDenseResidual (0.9084) trail baseline
  3. Plan variants (FineSpacing, LowOrder, NNLabel) β€” catastrophic failure on ACDC
  4. Generic inference postprocessing β€” temperature scaling, TTA, smoothing all ineffective
  5. FocalTversky loss β€” produced all-zero predictions
  6. GELU activation β€” marginal improvement, not worth the overhead
  7. TopK5 loss β€” too aggressive, slightly hurts
  8. L-plan models for 1000ep β€” insufficient capacity; the 3.18% gap of ComboLS_LateZ_DH shows smaller models can't match XL even with better training tricks

Generalization Status

Cross-dataset experiments on Dataset920 (KiTS2023) and Dataset923 (AMOS2022) have partial results. Some crashed and need reruns. This is tracked in launch_all_missing.py.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support