Burmese ASR
Collection
3 items โข Updated
This model is a fine-tuned version of facebook/wav2vec2-xls-r-300m for Automatic Speech Recognition (ASR) in Burmese .
my)The model was trained on a standardized, high-quality Burmese speech corpus:
\u1000-\u109F) with unified word segmentation and punctuation removal. The held-out test split evaluates zero-shot acoustic generalization across completely unseen voices.| Parameter | Value |
|---|---|
| Effective Batch Size | 32 (4 per device x 8 gradient accumulation) |
| Peak Learning Rate | 3.0e-4 |
| Learning Rate Scheduler | Linear |
| Warmup Steps | 200 |
| Total Steps | 2,000 (~6 Epochs) |
| Mixed Precision | BF16 |
| Gradient Checkpointing | Enabled (Input Require Grads) |
| Augmentation | None |
| Stage / Evaluation | Step | Train Loss | Val Loss | WER (%) | CER (%) | SER (%) | DER (%) | IER (%) | chrF |
|---|---|---|---|---|---|---|---|---|---|
| Baseline (Untrained) | 0 | - | - | 100.00 | 113.39 | 100.00 | 85.06 | 0.00 | 0.55 |
| Validation Set | 250 | 13.0478 | 1.1601 | 72.71 | 17.40 | 99.78 | 13.98 | 3.03 | 66.26 |
| Validation Set | 500 | 3.2769 | 0.4160 | 39.69 | 7.33 | 87.70 | 6.37 | 3.16 | 87.39 |
| Validation Set | 750 | 1.6812 | 0.3104 | 33.10 | 5.68 | 83.43 | 3.66 | 4.38 | 91.12 |
| Validation Set | 1000 | 1.1205 | 0.2926 | 29.82 | 4.68 | 78.35 | 5.50 | 2.43 | 92.75 |
| Validation Set | 1250 | 0.7550 | 0.2558 | 27.96 | 4.08 | 75.09 | 5.15 | 2.70 | 93.58 |
| Validation Set | 1500 | 0.3550 | 0.2688 | 25.96 | 3.81 | 71.17 | 4.16 | 3.21 | 94.43 |
| Validation Set | 1750 | 0.1300 | 0.3087 | 25.18 | 3.67 | 70.57 | 4.77 | 2.47 | 94.70 |
| Validation Set | 2000 | 0.0465 | 0.3125 | 24.25 | 3.53 | 69.13 | 3.77 | 3.07 | 94.97 |
| UNSEEN TEST (Final) | Final | - | - | 33.24 | 6.58 | 91.35 | 5.87 | 2.57 | 87.69 |
import torch
from transformers import pipeline
pipe = pipeline(
'automatic-speech-recognition',
model='thantzinphyo/Wav2Vec2-XLS-R-300M-ASR',
device='cuda:0' if torch.cuda.is_available() else 'cpu',
)
output = pipe('audio.wav')
print(output['text'])