Whisper-Tiny-Myanmar-UnFreezing

This model is a Full fine-tuned version of (https://huggingface.co/thantzinphyo/Whisper-Tiny-Myanmar-Partial-Freezing)) trained on thantzinphyo/burmese-speech-refined-openslr-80 Dataset.

Evaluation Results (Best Checkpoint: Epoch 7)

  • Best Validation WER: 54.8060%
  • Best Validation CER: 19.3867%
  • Validation Loss: 0.0970

Full 30-Epoch Training Results

Training Loss Epoch Step Validation Loss WER (%) CER (%) Status
No log 1.0 9 0.3240 66.7549 25.2793
5.8635 2.0 18 0.2111 64.0653 23.5311
3.9579 3.0 27 0.1359 60.9347 21.8387
2.2134 4.0 36 0.1084 59.4797 21.7605
1.4742 5.0 45 0.0987 57.6720 20.7775
1.0824 6.0 54 0.0955 56.8783 20.4088
0.8153 7.0 63 0.0970 54.8060 19.3867 Best Checkpoint
0.5734 8.0 72 0.1063 55.9083 20.2301
0.3814 9.0 81 0.1170 55.4674 20.0290
0.2371 10.0 90 0.1294 56.4374 20.3083
0.2371 11.0 99 0.1398 56.7460 20.2692
0.1447 12.0 108 0.1463 56.6578 20.1408
0.1012 13.0 117 0.1492 56.7019 20.1631
0.0755 14.0 126 0.1564 56.7460 20.6881
0.0588 15.0 135 0.1650 56.1287 20.2804
0.0488 16.0 144 0.1697 57.2751 20.0402
0.0431 17.0 153 0.1701 56.4815 20.1017
0.0314 18.0 162 0.1740 55.0265 19.9788
0.0249 19.0 171 0.1784 56.3492 19.9397
0.0214 20.0 180 0.1768 54.8501 19.6995
0.0214 21.0 189 0.1824 56.3051 19.9453
0.0167 22.0 198 0.1866 55.6878 19.7107
0.0107 23.0 207 0.1900 55.5556 19.5766
0.0087 24.0 216 0.1929 55.3792 19.3476
0.0078 25.0 225 0.1961 56.5697 19.8168
0.0055 26.0 234 0.1962 56.1287 19.7554
0.0035 27.0 243 0.1982 56.3492 20.0235
0.0022 28.0 252 0.2009 57.3192 20.1296
0.0016 29.0 261 0.2019 57.0988 20.0011
0.0017 30.0 270 0.2022 56.8342 20.0681

Hyperparameter Configuration

This model was trained using Full End-to-End Fine-Tuning (All Layers Unfrozen) on top of thantzinphyo/Whisper-Tiny-Myanmar-Partial-Freezing

Hyperparameter Value Description / Rationale
Base Model whisper-tiny-myanmar-phase1 Pre-aligned Burmese Whisper Tiny checkpoint
Training Mode Full Unfreeze (End-to-End) 100% of parameters (37.7M) updated simultaneously
Learning Rate 5e-05 (5e-5) Optimal sweet spot for fine-grained phonetic adaptation
LR Scheduler Linear with Warmup 100 warmup steps followed by linear decay
Per-Device Batch Size 16 Optimized for GPU VRAM throughput
Gradient Accumulation 8 Simulates a stable global batch size
Effective Batch Size 128 (16 × 8) Ensures low gradient variance and stable updates
Optimizer AdamW Weight decay of 0.01, $\beta_1 = 0.9$, $\beta_2 = 0.999$
Precision FP16 Mixed-precision for accelerated training and memory efficiency
Epochs 30 Best checkpoint selected at Epoch 7 (load_best_model_at_end)

Downloads last month
-
Safetensors
Model size
37.8M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thantzinphyo/Whisper-Tiny-Myanmar-UnFreezing

Finetuned
(1)
this model

Dataset used to train thantzinphyo/Whisper-Tiny-Myanmar-UnFreezing

Evaluation results

  • Validation WER on Burmese Speech Refined OpenSLR-80
    self-reported
    54.806
  • Validation CER on Burmese Speech Refined OpenSLR-80
    self-reported
    19.387