Whisper-Base-ASR ( Burmese )

This model is a fine-tuned version of openai/whisper-base for Automatic Speech Recognition (ASR) in Burmese .

Model Overview

  • Architecture: Whisper (Base)
  • Parameters: ~74M
  • Task: Automatic Speech Recognition (ASR)
  • Language: Burmese (my)
  • Sampling Rate: 16,000 Hz (Mono)

Dataset Details

The model was trained on a standardized, high-quality Burmese speech corpus:

  • Total Duration: ~22 Hours of audio
  • Total Utterances: 24,560 WAV files (16 kHz, 16-bit PCM, Mono)
  • Total Speakers: 13 Synthetic Burmese Speakers
  • Speakers: ပီယ ၊ ဝါစာ ၊ သီရိ ၊ ဒီပ ၊ အက္ခရာ ၊ သဒ္ဒါ ၊ သရ ၊ တာရာ ၊ ကဝိ ၊ သုတ ၊ ပသာဒ ၊ ဂီတ ၊ နန္ဒ
  • Training & Validation: 11 speakers (20,699 train utterances, 2,300 validation utterances)
  • Test Set (Held-Out): 2 unseen speakers ( ဂီတ & နန္ဒ , 1,561 utterances )
  • Text Preprocessing: Standardized Myanmar Unicode (\u1000-\u109F) with unified word segmentation and punctuation removal. The held-out test split evaluates zero-shot acoustic generalization across completely unseen voices.

Training Configuration

Parameter Value
Effective Batch Size 64 (16 per device x 4 gradient accumulation)
Peak Learning Rate 1.0e-4
Learning Rate Scheduler Cosine
Warmup Steps 200
Total Steps 2,000 (~12.37 Epochs)
Mixed Precision BF16
Weight Decay 0.01
Gradient Checkpointing Enabled
Augmentation None

Evaluation Results

Stage / Evaluation Step Train Loss Val Loss WER (%) CER (%) SER (%) DER (%) IER (%) chrF
Baseline (Untrained) 0 - - 753.85 437.94 100.00 7.69 653.85 0.00
Validation Set 250 0.6374 0.0725 47.99 10.33 93.57 10.36 2.65 81.45
Validation Set 500 0.2351 0.0453 37.15 6.92 87.52 6.12 2.81 86.66
Validation Set 750 0.1369 0.0417 31.71 5.70 81.61 3.60 4.19 89.08
Validation Set 1000 0.0482 0.0416 28.87 5.00 76.74 4.22 3.29 91.07
Validation Set 1250 0.0252 0.0435 28.18 4.79 76.74 3.72 3.31 91.23
Validation Set 1500 0.0039 0.0476 28.06 4.69 78.48 3.81 2.94 91.40
Validation Set 1750 0.0007 0.0502 26.91 4.47 76.30 3.26 3.44 91.79
Validation Set 2000 0.0004 0.0507 26.55 4.43 75.65 3.31 3.37 91.94
UNSEEN TEST (Final) Final - - 36.29 8.55 94.30 5.40 3.22 83.71

Usage

import torch
from transformers import pipeline

pipe = pipeline(
    'automatic-speech-recognition',
    model='thantzinphyo/Whisper-Base-ASR',
    device='cuda:0' if torch.cuda.is_available() else 'cpu',
)

output = pipe('audio.wav', generate_kwargs={'language': 'burmese', 'task': 'transcribe'})
print(output['text'])

License

Apache-2.0

Downloads last month
76
Safetensors
Model size
72.6M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thantzinphyo/Whisper-Base-ASR

Finetuned
(763)
this model

Dataset used to train thantzinphyo/Whisper-Base-ASR

Collection including thantzinphyo/Whisper-Base-ASR

Evaluation results

  • Validation CER on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    4.430
  • Validation WER on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    26.550
  • Validation chrF on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    91.940
  • Test CER (Unseen Speakers) on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    8.550
  • Test WER (Unseen Speakers) on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    36.290
  • Test chrF (Unseen Speakers) on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    83.710