Whisper-Tiny-Myanmar-Partial-Freezing

Fine-tuned version of openai/whisper-tiny on SLR80: Crowdsourced High-Quality Burmese Speech Dataset ( https://www.openslr.org/80 ).

Best Model Selection Notice:

This repository contains the best checkpoint (Step 1,000), selected for its lowest CER of 19.60%.


How to Use (Python Inference)

import torch
from transformers import pipeline

pipe = pipeline(
    "automatic-speech-recognition",
    model="thantzinphyo/Whisper-Tiny-Myanmar-Partial-Freezing",
    device=0 if torch.cuda.is_available() else -1
)

result = pipe("your_burmese_audio.wav", generate_kwargs={"language": "my", "task": "transcribe"})
print("Transcribed Text:", result["text"]) # တချို့ α€€α€œα€Šα€Ία€Έ α€•α€Όα€±α€¬α€α€šα€Ί α€™α€„α€Ία€Ήα€‚α€œα€¬α€†α€±α€¬α€„α€Ία€α€šα€Ί ဆိုတာ ထိမ်ထောင်ကျတာ α€‘α€±α€¬α€„α€Ία€€α€»α€α€šα€Ί ပေါ့

Complete Training & Evaluation History (Steps 0 to 1,500)

Evaluated on the isolated validation set (252 clips / ~0.42 hrs) and formatted in Percentage (%):

Step Training Loss Validation Loss Character Error Rate (CER) Word Error Rate (WER) Checkpoint Status
0 (Baseline) β€” β€” 230.57% 213.09% Pretrained Base (Hallucination)
100 17.2880 2.1209 78.09% 100.04% Script Recognition
200 13.1279 1.6428 33.08% 81.22% Phonetic Alignment
300 11.7845 1.6119 27.40% 73.99% Syntax Formation
400 11.4952 1.6188 24.47% 69.96% Sub-70% WER Barrier Broken
500 11.4360 1.6261 24.06% 70.84% Vocabulary Consolidation
600 11.3406 1.6402 20.95% 66.37% Particle Tuning
700 11.3297 1.6517 22.98% 69.09% Regularization
800 11.3435 1.6443 21.23% 67.29% Consonant Conjunct Refinement
900 11.3192 1.6535 20.37% 66.07% Stable Convergence
1,000 11.2989 1.6744 19.60% 65.28% Checkpoint (Published)
1,100 11.2955 1.6819 19.80% 64.97% Lowest WER Record
1,200 11.2928 1.6945 19.71% 65.15% High Precision
1,300 11.2914 1.7026 19.72% 65.19% Cosine Learning Decay
1,400 11.2909 1.7060 19.94% 65.46% Asymptotic Convergence
1,500 11.2908 1.7066 19.88% 65.37% Final Step Completion

Initial Training Configuration

  • Base Architecture: openai/whisper-tiny (37.8M parameters)
  • Training Strategy: Phase 1 Partial Encoder Freeze β€” Layers 0–1 frozen, while Layers 2–3 and the Decoder were trainable (33.1M parameters / 87.66%)
  • Effective Batch Size: 64 (16 per device Γ— 4 gradient accumulation)
  • Learning Rate: 1.5e-4 with a Cosine Scheduler and 150 warmup steps
  • Dataset: 2,277 training clips and 252 validation clips using 100% OpenSLR-80 NFC-canonical transcriptions
Downloads last month
70
Safetensors
Model size
37.8M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for thantzinphyo/Whisper-Tiny-Myanmar-Partial-Freezing

Finetunes
1 model