thantzinphyo/Burmese-Daily-Dialogue-Corpus
Viewer • Updated • 24.6k • 120 • 1
This model is a fine-tuned version of facebook/mms-300m for Automatic Speech Recognition (ASR) in Burmese .
my)The model was trained on a standardized, high-quality Burmese speech corpus:
| Parameter | Value |
|---|---|
| Effective Batch Size | 64 (8 per device x 4 gradient accumulation) |
| Base Architecture | facebook/mms-300m ( Wav2Vec 2.0 with 24 Transformer layers, 1024 hidden size, 16 attention heads ) |
| Acoustic Feature Extractor | 16 kHz 1D raw waveform input (CNN encoder frozen during training) |
| Tokenization | Burmese Unicode character CTC vocabulary |
| Peak Learning Rate | 3.0e-4 |
| Learning Rate Scheduler | Linear |
| Warmup Steps | 200 |
| Total Steps | 2,000 (~6 Epochs) |
| Gradient Checkpointing | Enabled (Input Require Grads) |
| Augmentation | None |
| Stage / Evaluation | Step | Train Loss | Val Loss | WER (%) | CER (%) | SER (%) | DER (%) | IER (%) | chrF |
|---|---|---|---|---|---|---|---|---|---|
| Baseline (Untrained) | 0 | - | - | 100.00 | 121.74 | 100.00 | 84.62 | 0.00 | 1.30 |
| Validation Set | 250 | 6.4141 | 1.0876 | 67.66 | 14.87 | 99.17 | 10.48 | 4.22 | 69.03 |
| Validation Set | 500 | 1.5876 | 0.4096 | 37.75 | 6.34 | 85.74 | 6.62 | 2.86 | 87.67 |
| Validation Set | 750 | 0.8111 | 0.2971 | 28.93 | 4.44 | 74.65 | 3.76 | 4.08 | 91.69 |
| Validation Set | 1000 | 0.5157 | 0.2717 | 26.27 | 3.84 | 70.91 | 3.76 | 3.59 | 93.22 |
| Validation Set | 1250 | 0.3567 | 0.2459 | 24.87 | 3.56 | 67.22 | 4.20 | 2.94 | 94.06 |
| Validation Set | 1500 | 0.1711 | 0.2656 | 23.40 | 3.25 | 64.09 | 2.83 | 4.11 | 94.56 |
| Validation Set | 1750 | 0.0564 | 0.2951 | 21.23 | 2.93 | 60.22 | 3.20 | 3.34 | 95.22 |
| Validation Set | 2000 | 0.0208 | 0.2972 | 20.66 | 2.85 | 59.35 | 3.12 | 3.15 | 95.39 |
| UNSEEN TEST (Final) | Final | - | - | 32.64 | 6.66 | 90.01 | 5.33 | 2.60 | 87.22 |
import torch
from transformers import pipeline
pipe = pipeline(
'automatic-speech-recognition',
model='thantzinphyo/Meta-MMS-300M-ASR',
device='cuda:0' if torch.cuda.is_available() else 'cpu',
)
output = pipe('audio.wav')
print(output['text'])
Base model
facebook/mms-300m