Whisper-Small-ASR ( Burmese )

This model is a fine-tuned version of openai/whisper-small for Automatic Speech Recognition (ASR) in Burmese .

Model Overview

  • Architecture: Whisper Small
  • Parameters: ~244M
  • Task: Automatic Speech Recognition (ASR)
  • Language: Burmese (my)
  • Sampling Rate: 16,000 Hz (Mono)

Dataset Details

The model was trained on a standardized, high-quality Burmese speech corpus:

  • Total Duration: ~22 Hours of audio
  • Total Utterances: 24,560 WAV files (16 kHz, 16-bit PCM, Mono)
  • Total Speakers: 13 Synthetic Burmese Speakers
  • Speakers: ပီယ ၊ ဝါစာ ၊ သီရိ ၊ ဒီပ ၊ အက္ခရာ ၊ သဒ္ဒါ ၊ သရ ၊ တာရာ ၊ ကဝိ ၊ သုတ ၊ ပသာဒ ၊ ဂီတ ၊ နန္ဒ
  • Training & Validation: 11 speakers (20,699 train utterances, 2,300 validation utterances)
  • Test Set (Held-Out): 2 unseen speakers ( ဂီတ & နန္ဒ , 1,561 utterances )
  • Text Preprocessing: Standardized Myanmar Unicode (\u1000-\u109F) with unified word segmentation and punctuation removal. The held-out test split evaluates zero-shot acoustic generalization across completely unseen voices.

Training Configuration

Parameter Value
Effective Batch Size 64 (8 per device x 4 gradient accumulation x 2 GPUs)
Peak Learning Rate 2.5e-5
Learning Rate Scheduler Cosine
Warmup Steps 200
Total Steps 2,000 (~6 Epochs)
Mixed Precision FP16
Gradient Checkpointing Enabled
Augmentation None

Evaluation Results

Stage / Evaluation Step Train Loss Val Loss WER (%) CER (%) SER (%) DER (%) IER (%) chrF
Baseline (Untrained) 0 - - 610.77 441.11 100.00 44.62 510.77 0.00
Validation Set 250 0.9354 0.1031 60.56 14.25 97.70 18.50 1.39 74.65
Validation Set 500 0.3638 0.0493 33.98 6.46 82.16 4.85 4.22 88.13
Validation Set 750 0.1874 0.0387 28.06 5.14 71.61 4.94 3.03 91.09
Validation Set 1000 0.0869 0.0338 23.82 3.90 65.54 3.88 3.27 93.54
Validation Set 1250 0.0842 0.0314 23.44 3.70 64.97 5.46 2.16 94.00
Validation Set 1500 0.0319 0.0339 21.29 3.37 60.33 3.31 3.47 94.58
Validation Set 1750 0.0080 0.0372 20.96 3.25 58.98 3.87 2.91 94.80
Validation Set 2000 0.0043 0.0375 20.70 3.21 58.59 3.87 2.78 94.86
UNSEEN TEST (Final) Final - - 28.22 5.74 84.38 6.34 2.47 89.93

Usage

import torch
from transformers import pipeline

pipe = pipeline(
    'automatic-speech-recognition',
    model='thantzinphyo/Whisper-Small-ASR',
    device='cuda:0' if torch.cuda.is_available() else 'cpu',
)

output = pipe('audio.wav', generate_kwargs={"language": "my", "task": "transcribe"})
print(output['text']) # ငါ့ အချစ်ရေးကျတော့ ဘာလို့ မကောင်းတာလဲ ဆိုတာ ဘယ်သူမှ မပြောနိုင်ဘူး

License

Apache-2.0

Downloads last month
31
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thantzinphyo/Whisper-Small-ASR

Finetuned
(3765)
this model

Dataset used to train thantzinphyo/Whisper-Small-ASR

Collection including thantzinphyo/Whisper-Small-ASR

Evaluation results