Whisper-Tiny-Myanmar-Partial-Freezing
Fine-tuned version of openai/whisper-tiny on SLR80: Crowdsourced High-Quality Burmese Speech Dataset ( https://www.openslr.org/80 ).
Best Model Selection Notice:
This repository contains the best checkpoint (Step 1,000), selected for its lowest CER of 19.60%.
How to Use (Python Inference)
import torch
from transformers import pipeline
pipe = pipeline(
"automatic-speech-recognition",
model="thantzinphyo/Whisper-Tiny-Myanmar-Partial-Freezing",
device=0 if torch.cuda.is_available() else -1
)
result = pipe("your_burmese_audio.wav", generate_kwargs={"language": "my", "task": "transcribe"})
print("Transcribed Text:", result["text"]) # ααα»αα―α· ααααΊαΈ ααΌα±α¬αααΊ αααΊαΉααα¬αα±α¬ααΊαααΊ ααα―αα¬ α‘αααΊαα±α¬ααΊαα»αα¬ αα±α¬ααΊαα»αααΊ αα±α«α·
Complete Training & Evaluation History (Steps 0 to 1,500)
Evaluated on the isolated validation set (252 clips / ~0.42 hrs) and formatted in Percentage (%):
| Step | Training Loss | Validation Loss | Character Error Rate (CER) | Word Error Rate (WER) | Checkpoint Status |
|---|---|---|---|---|---|
| 0 (Baseline) | β | β | 230.57% | 213.09% | Pretrained Base (Hallucination) |
| 100 | 17.2880 | 2.1209 | 78.09% | 100.04% | Script Recognition |
| 200 | 13.1279 | 1.6428 | 33.08% | 81.22% | Phonetic Alignment |
| 300 | 11.7845 | 1.6119 | 27.40% | 73.99% | Syntax Formation |
| 400 | 11.4952 | 1.6188 | 24.47% | 69.96% | Sub-70% WER Barrier Broken |
| 500 | 11.4360 | 1.6261 | 24.06% | 70.84% | Vocabulary Consolidation |
| 600 | 11.3406 | 1.6402 | 20.95% | 66.37% | Particle Tuning |
| 700 | 11.3297 | 1.6517 | 22.98% | 69.09% | Regularization |
| 800 | 11.3435 | 1.6443 | 21.23% | 67.29% | Consonant Conjunct Refinement |
| 900 | 11.3192 | 1.6535 | 20.37% | 66.07% | Stable Convergence |
| 1,000 | 11.2989 | 1.6744 | 19.60% | 65.28% | Checkpoint (Published) |
| 1,100 | 11.2955 | 1.6819 | 19.80% | 64.97% | Lowest WER Record |
| 1,200 | 11.2928 | 1.6945 | 19.71% | 65.15% | High Precision |
| 1,300 | 11.2914 | 1.7026 | 19.72% | 65.19% | Cosine Learning Decay |
| 1,400 | 11.2909 | 1.7060 | 19.94% | 65.46% | Asymptotic Convergence |
| 1,500 | 11.2908 | 1.7066 | 19.88% | 65.37% | Final Step Completion |
Initial Training Configuration
- Base Architecture:
openai/whisper-tiny(37.8M parameters) - Training Strategy: Phase 1 Partial Encoder Freeze β Layers 0β1 frozen, while Layers 2β3 and the Decoder were trainable (33.1M parameters / 87.66%)
- Effective Batch Size: 64 (16 per device Γ 4 gradient accumulation)
- Learning Rate:
1.5e-4with a Cosine Scheduler and 150 warmup steps - Dataset: 2,277 training clips and 252 validation clips using 100% OpenSLR-80 NFC-canonical transcriptions
- Downloads last month
- 70