π²π² Whisper-Tiny Myanmar ASR (Phase 1)
Fine-tuned version of openai/whisper-tiny on OpenSLR-80 (Crowdsourced Burmese Speech) dataset.
π Best Model Selection Notice:
This repository hosts the best-performing checkpoint (Step 1,000) selected automatically from 1,500 training steps based on the lowest Character Error Rate (CER = 19.60%), effectively preventing late-stage overfitting.
π Complete Training & Evaluation History (Steps 0 to 1,500)
Evaluated on the isolated validation set (253 clips / ~0.42 hrs) and formatted in Percentage (%):
| Step | Training Loss | Validation Loss | Character Error Rate (CER) | Word Error Rate (WER) | Checkpoint Status |
|---|---|---|---|---|---|
| 0 (Baseline) | β | β | 230.57% | 213.09% | Pretrained Base (Hallucination) |
| 100 | 17.2880 | 2.1209 | 78.09% | 100.04% | Script Recognition |
| 200 | 13.1279 | 1.6428 | 33.08% | 81.22% | Phonetic Alignment |
| 300 | 11.7845 | 1.6119 | 27.40% | 73.99% | Syntax Formation |
| 400 | 11.4952 | 1.6188 | 24.47% | 69.96% | Sub-70% WER Barrier Broken |
| 500 | 11.4360 | 1.6261 | 24.06% | 70.84% | Vocabulary Consolidation |
| 600 | 11.3406 | 1.6402 | 20.95% | 66.37% | Particle Tuning |
| 700 | 11.3297 | 1.6517 | 22.98% | 69.09% | Regularization |
| 800 | 11.3435 | 1.6443 | 21.23% | 67.29% | Consonant Conjunct Refinement |
| 900 | 11.3192 | 1.6535 | 20.37% | 66.07% | Stable Convergence |
| 1,000 | 11.2989 | 1.6744 | 19.60% π― | 65.28% | π BEST CHECKPOINT (Published) |
| 1,100 | 11.2955 | 1.6819 | 19.80% | 64.97% π₯ | Lowest WER Record |
| 1,200 | 11.2928 | 1.6945 | 19.71% | 65.15% | High Precision |
| 1,300 | 11.2914 | 1.7026 | 19.72% | 65.19% | Cosine Learning Decay |
| 1,400 | 11.2909 | 1.7060 | 19.94% | 65.46% | Asymptotic Convergence |
| 1,500 | 11.2908 | 1.7066 | 19.88% | 65.37% | Final Step Completion |
βοΈ Hyperparameters & Training Specifications
- Base Architecture:
openai/whisper-tiny(37.8M parameters) - Training Strategy: Phase 1 Partial Encoder Freeze (Layers 0-1 frozen, Layers 2-3 + Decoder trainable: 33.1M params / 87.66%)
- Effective Batch Size: 64 (16 per device x 4 gradient accumulation)
- Learning Rate:
1.5e-4with Cosine Scheduler & 150 warmup steps - Dataset: 2,277 train clips / 253 validation clips (100% Exact OpenSLR-80 NFC canonical transcription)
- Downloads last month
- 41