Meta-MMS-300M-ASR ( Burmese )

This model is a fine-tuned version of facebook/mms-300m for Automatic Speech Recognition (ASR) in Burmese .

Model Overview

  • Architecture: Meta MMS (300M)
  • Parameters: ~317M
  • Task: Automatic Speech Recognition (ASR)
  • Language: Burmese (my)
  • Sampling Rate: 16,000 Hz (Mono)
  • Tokenizer / Vocab: 64 Burmese Character Tokens (CTC Head)

Dataset Details

The model was trained on a standardized, high-quality Burmese speech corpus:

  • Total Duration: ~22 Hours of audio
  • Total Utterances: 24,560 WAV files (16 kHz, 16-bit PCM, Mono)
  • Total Speakers: 13 Synthetic Burmese Speakers
  • Speakers: ပီယ ၊ ဝါစာ ၊ သီရိ ၊ ဒီပ ၊ အက္ခရာ ၊ သဒ္ဒါ ၊ သရ ၊ တာရာ ၊ ကဝိ ၊ သုတ ၊ ပသာဒ ၊ ဂီတ ၊ နန္ဒ
  • Training & Validation: 11 speakers (20,699 train utterances, 2,300 validation utterances)
  • Test Set (Held-Out): 2 unseen speakers ( ဂီတ & နန္ဒ , 1,561 utterances )
  • Text Preprocessing: Standardized Myanmar Unicode with unified word segmentation and punctuation removal. The held-out test split evaluates zero-shot acoustic generalization across completely unseen voices.

Training Configuration

Parameter Value
Effective Batch Size 64 (8 per device x 4 gradient accumulation)
Base Architecture facebook/mms-300m ( Wav2Vec 2.0 with 24 Transformer layers, 1024 hidden size, 16 attention heads )
Acoustic Feature Extractor 16 kHz 1D raw waveform input (CNN encoder frozen during training)
Tokenization Burmese Unicode character CTC vocabulary
Peak Learning Rate 3.0e-4
Learning Rate Scheduler Linear
Warmup Steps 200
Total Steps 2,000 (~6 Epochs)
Gradient Checkpointing Enabled (Input Require Grads)
Augmentation None

Evaluation Results

Stage / Evaluation Step Train Loss Val Loss WER (%) CER (%) SER (%) DER (%) IER (%) chrF
Baseline (Untrained) 0 - - 100.00 121.74 100.00 84.62 0.00 1.30
Validation Set 250 6.4141 1.0876 67.66 14.87 99.17 10.48 4.22 69.03
Validation Set 500 1.5876 0.4096 37.75 6.34 85.74 6.62 2.86 87.67
Validation Set 750 0.8111 0.2971 28.93 4.44 74.65 3.76 4.08 91.69
Validation Set 1000 0.5157 0.2717 26.27 3.84 70.91 3.76 3.59 93.22
Validation Set 1250 0.3567 0.2459 24.87 3.56 67.22 4.20 2.94 94.06
Validation Set 1500 0.1711 0.2656 23.40 3.25 64.09 2.83 4.11 94.56
Validation Set 1750 0.0564 0.2951 21.23 2.93 60.22 3.20 3.34 95.22
Validation Set 2000 0.0208 0.2972 20.66 2.85 59.35 3.12 3.15 95.39
UNSEEN TEST (Final) Final - - 32.64 6.66 90.01 5.33 2.60 87.22

Usage

import torch
from transformers import pipeline

pipe = pipeline(
    'automatic-speech-recognition',
    model='thantzinphyo/Meta-MMS-300M-ASR',
    device='cuda:0' if torch.cuda.is_available() else 'cpu',
)

output = pipe('audio.wav')
print(output['text'])

References & Citations

Downloads last month
37
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thantzinphyo/Meta-MMS-300M-ASR

Finetuned
(61)
this model

Dataset used to train thantzinphyo/Meta-MMS-300M-ASR

Collection including thantzinphyo/Meta-MMS-300M-ASR

Papers for thantzinphyo/Meta-MMS-300M-ASR

Evaluation results

  • Validation CER on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    2.850
  • Validation WER on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    20.660
  • Validation chrF on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    95.390
  • Test CER (Unseen Speakers) on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    6.660
  • Test WER (Unseen Speakers) on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    32.640
  • Test chrF (Unseen Speakers) on Burmese Daily Dialogue Corpus (BDDC)
    self-reported
    87.220