whisper-small-santali-coded
A full fine-tune of openai/whisper-small for automatic speech recognition on Santali, transcribed into Sanlish β a Latin-script romanization scheme for Santali speech.
Model description
Unlike partial-freeze / adapter-style approaches, every parameter of the base Whisper-small checkpoint (encoder, decoder, and output head) was fine-tuned. No language tag was pinned during tokenization or generation (generation_config.language = None), so the model was not steered toward Bengali or any other Whisper-supported language β it learned Sanlish surface forms purely from the fine-tuning data.
- Base model:
openai/whisper-small(~244M params) - Fine-tuning strategy: full fine-tune, no frozen layers
- Task: speech transcription (audio β Sanlish text)
- Language tag: none pinned (language-agnostic decoding)
Training data
- Source: custom Santali audio + Sanlish transliteration dataset
- Splits: train / validation / test (
santali_train.csv,santali_val.csv,santali_test.csv) - Format:
audio_path(16kHz audio) βsanlish(target transliteration text)
Training procedure
| Hyperparameter | Value |
|---|---|
| Base checkpoint | openai/whisper-small |
| Learning rate | 2e-5 |
| Weight decay | 0.1 |
| Batch size (train) | 16 |
| Epochs (planned / completed) | 20 planned, 17 completed (training interrupted by a runtime disconnect) |
| Warmup steps | computed dynamically as 10% of total training steps |
| Eval/save strategy | every epoch |
| Early stopping | patience 5 (on validation WER) β not yet triggered when training stopped |
| Metric for best model | WER (lower is better) |
Evaluation results
Per-epoch validation metrics (normalized WER/CER β lowercased, punctuation-stripped):
| Epoch | Train Loss | Val Loss | WER (%) | CER (%) |
|---|---|---|---|---|
| 1 | 1.8145 | 0.6641 | 54.51 | 11.61 |
| 3 | 0.1810 | 0.4263 | 36.37 | 6.60 |
| 8 | 0.0142 | 0.4845 | 33.30 | 5.92 |
| 10 | 0.0068 | 0.4903 | 31.50 | 5.58 |
| 14 (best WER) | 0.0006 | 0.5197 | 31.07 | 5.39 |
17 (checkpoint on main) |
0.0003 | 0.5266 | 31.18 | 5.41 |
Note: validation loss bottoms out around epoch 3 and rises afterward while WER/CER keep improving slightly β a sign of overfitting on the loss objective that hasn't yet hurt transcription accuracy. Epoch 14 had the lowest validation WER of the run; epoch 17 (the last checkpoint pushed before the training runtime disconnected) is marginally behind it and is what's currently on main. Earlier epoch checkpoints are available in this repo's commit history.
Test set: 28.82% WER, 4.89% CER β evaluated on the held-out test split (193 examples), using the epoch-17 checkpoint currently on main.
Limitations
- Training did not complete its full 20-epoch schedule (interrupted at epoch 17), and early stopping had not yet triggered.
- Fine-tuned on a relatively small custom dataset for a low-resource language variant; may not generalize to speakers, dialects, or recording conditions outside the training distribution.
- No language tag is pinned, so downstream users should explicitly pass
task="transcribe"at inference and not rely on Whisper's automatic language detection for this checkpoint.
Usage
from transformers import WhisperForConditionalGeneration, WhisperProcessor
model = WhisperForConditionalGeneration.from_pretrained("thunderboltc/whisper-small-santali-coded")
processor = WhisperProcessor.from_pretrained("thunderboltc/whisper-small-santali-coded", task="transcribe")
model.generation_config.language = None
model.generation_config.task = "transcribe"
# input_features = processor(audio_array, sampling_rate=16000, return_tensors="pt").input_features
# predicted_ids = model.generate(input_features)
# transcription = processor.batch_decode(predicted_ids, skip_special_tokens=True)
- Downloads last month
- 158
Model tree for thunderboltc/whisper-small-santali-coded
Base model
openai/whisper-smallEvaluation results
- Validation WER (epoch 17)self-reported31.180
- Validation CER (epoch 17)self-reported5.410
- Test WER (epoch 17)self-reported28.820
- Test CER (epoch 17)self-reported4.890