Whisper-Tiny-ASR ( Burmese )

This model is a fine-tuned version of openai/whisper-tiny for Automatic Speech Recognition (ASR) in Burmese .

Model Overview

  • Architecture: Whisper (Tiny)
  • Parameters: ~39M
  • Task: Automatic Speech Recognition (ASR)
  • Language: Burmese (my)
  • Sampling Rate: 16,000 Hz (Mono)

Dataset Details

The model was trained on a standardized, high-quality Burmese speech corpus:

  • Total Duration: ~22 Hours of audio
  • Total Utterances: 24,560 WAV files (16 kHz, 16-bit PCM, Mono)
  • Total Speakers: 13 Synthetic Burmese Speakers
  • Speakers: แ€•แ€ฎแ€š แŠ แ€แ€ซแ€…แ€ฌ แŠ แ€žแ€ฎแ€›แ€ญ แŠ แ€’แ€ฎแ€• แŠ แ€กแ€€แ€นแ€แ€›แ€ฌ แŠ แ€žแ€’แ€นแ€’แ€ซ แŠ แ€žแ€› แŠ แ€แ€ฌแ€›แ€ฌ แŠ แ€€แ€แ€ญ แŠ แ€žแ€ฏแ€ แŠ แ€•แ€žแ€ฌแ€’ แŠ แ€‚แ€ฎแ€ แŠ แ€”แ€”แ€นแ€’
  • Training & Validation: 11 speakers (20,699 train utterances, 2,300 validation utterances)
  • Test Set (Held-Out): 2 unseen speakers ( แ€‚แ€ฎแ€ & แ€”แ€”แ€นแ€’ , 1,561 utterances )
  • Text Preprocessing: Standardized Myanmar Unicode (\u1000-\u109F) with unified word segmentation and punctuation removal. The held-out test split evaluates zero-shot acoustic generalization across completely unseen voices.

Training Configuration

Parameter Value
Effective Batch Size 64 (16 per device x 4 gradient accumulation)
Peak Learning Rate 1.5e-4
Learning Rate Scheduler Cosine
Warmup Steps 150
Total Steps 1,500 (~9.26 Epochs)
Mixed Precision BF16
Weight Decay 0.01
Gradient Checkpointing Enabled
Augmentation None

Evaluation Results

Stage / Evaluation Step Train Loss Val Loss WER (%) CER (%) SER (%) DER (%) IER (%) chrF
Baseline (Untrained) 0 - - 104.62 130.83 100.00 41.54 4.62 0.00
Validation Set 250 0.9590 0.1098 60.79 19.93 97.61 10.94 5.76 69.02
Validation Set 500 0.3188 0.0551 38.17 8.25 84.78 8.45 2.90 86.15
Validation Set 750 0.1719 0.0428 30.01 5.98 75.13 5.39 3.62 90.26
Validation Set 1000 0.0447 0.0442 26.92 5.17 69.13 5.18 3.08 91.85
Validation Set 1250 0.0119 0.0492 25.66 4.57 66.83 4.57 3.45 92.78
Validation Set 1500 0.0026 0.0508 25.21 4.47 65.52 4.47 3.42 92.91
UNSEEN TEST (Final) Final - - 42.29 11.69 94.17 7.99 3.47 79.82

Usage

import torch
from transformers import pipeline

pipe = pipeline(
    "automatic-speech-recognition",
    model="thantzinphyo/Whisper-Tiny-ASR",
    device="cuda:0" if torch.cuda.is_available() else "cpu",
)

output = pipe("audio.wav", generate_kwargs={"language": "burmese", "task": "transcribe"})
print(output["text"])

License

Apache-2.0

Downloads last month
30
Safetensors
Model size
37.8M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Collection including thantzinphyo/Whisper-Tiny-ASR