πŸ‡²πŸ‡² Whisper-Tiny Myanmar ASR (Phase 1)

Fine-tuned version of openai/whisper-tiny on OpenSLR-80 (Crowdsourced Burmese Speech) dataset.

πŸ† Best Model Selection Notice:
This repository hosts the best-performing checkpoint (Step 1,000) selected automatically from 1,500 training steps based on the lowest Character Error Rate (CER = 19.60%), effectively preventing late-stage overfitting.


πŸ“Š Complete Training & Evaluation History (Steps 0 to 1,500)

Evaluated on the isolated validation set (253 clips / ~0.42 hrs) and formatted in Percentage (%):

Step Training Loss Validation Loss Character Error Rate (CER) Word Error Rate (WER) Checkpoint Status
0 (Baseline) β€” β€” 230.57% 213.09% Pretrained Base (Hallucination)
100 17.2880 2.1209 78.09% 100.04% Script Recognition
200 13.1279 1.6428 33.08% 81.22% Phonetic Alignment
300 11.7845 1.6119 27.40% 73.99% Syntax Formation
400 11.4952 1.6188 24.47% 69.96% Sub-70% WER Barrier Broken
500 11.4360 1.6261 24.06% 70.84% Vocabulary Consolidation
600 11.3406 1.6402 20.95% 66.37% Particle Tuning
700 11.3297 1.6517 22.98% 69.09% Regularization
800 11.3435 1.6443 21.23% 67.29% Consonant Conjunct Refinement
900 11.3192 1.6535 20.37% 66.07% Stable Convergence
1,000 11.2989 1.6744 19.60% 🎯 65.28% πŸ† BEST CHECKPOINT (Published)
1,100 11.2955 1.6819 19.80% 64.97% πŸ”₯ Lowest WER Record
1,200 11.2928 1.6945 19.71% 65.15% High Precision
1,300 11.2914 1.7026 19.72% 65.19% Cosine Learning Decay
1,400 11.2909 1.7060 19.94% 65.46% Asymptotic Convergence
1,500 11.2908 1.7066 19.88% 65.37% Final Step Completion

βš™οΈ Hyperparameters & Training Specifications

  • Base Architecture: openai/whisper-tiny (37.8M parameters)
  • Training Strategy: Phase 1 Partial Encoder Freeze (Layers 0-1 frozen, Layers 2-3 + Decoder trainable: 33.1M params / 87.66%)
  • Effective Batch Size: 64 (16 per device x 4 gradient accumulation)
  • Learning Rate: 1.5e-4 with Cosine Scheduler & 150 warmup steps
  • Dataset: 2,277 train clips / 253 validation clips (100% Exact OpenSLR-80 NFC canonical transcription)
Downloads last month
41
Safetensors
Model size
37.8M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for thantzinphyo/whisper-tiny-myanmar-phase1

Finetunes
1 model