Instructions to use h-gajdov/full_moonshine_base_all_data_wup1_ep3_lr2e-4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use h-gajdov/full_moonshine_base_all_data_wup1_ep3_lr2e-4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="h-gajdov/full_moonshine_base_all_data_wup1_ep3_lr2e-4")# Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("h-gajdov/full_moonshine_base_all_data_wup1_ep3_lr2e-4") model = AutoModelForSpeechSeq2Seq.from_pretrained("h-gajdov/full_moonshine_base_all_data_wup1_ep3_lr2e-4", device_map="auto") - Notebooks
- Google Colab
- Kaggle
full_moonshine_base_all_data_wup1_ep3_lr2e-4
Full fine-tune of UsefulSensors/moonshine-base (61.5M parameters) for Macedonian speech-to-text, trained on the full ~2.99M-clip corpus of the study.
Moonshine is a raw-waveform encoder-decoder: it takes input_values rather than
Whisper's fixed 30 s log-mel window, so its cost scales with actual audio length
instead of being flat per clip. That is what makes it viable on-device, and it is
why its real-time factor barely moves with utterance length where Whisper's does.
This is one of eight runs sweeping peak learning rate and warmup ratio over an
otherwise identical recipe -- the same sweep already run on moonshine-tiny, at
2.2x the parameters. All eight are in
this collection.
Evaluation
Held-out test set rachno_provereno_od_yt (44 manually verified clips, 0.32 h),
never seen in training. CPU, float32, batch 1, greedy (1 beam), no silence
padding, no length floor.
| metric | value |
|---|---|
| WER | 8.34 |
| CER | 3.01 |
| SER | 97.73 |
| Parameters | 61.5 M |
| Weights | 246 MB |
| Latency | 1930 ms/utterance |
| Throughput | 14x real-time |
| RTF | 0.0732 |
| Checkpoint | step 215,000 of 284,226 (epoch 2.27 of 3) |
SER is exact-match over whole utterances. These test clips average 62 words, so at this WER an exactly-correct utterance is vanishingly unlikely -- read WER/CER.
Best WER seen during training (same held-out set): 8.20 at step 215,000. That comes from the trainer's own in-loop generation -- batched, greedy -- so it is a training signal rather than a number to compare across projects; on these eight runs it lands within 0.25 WER of the table above.
This checkpoint is not the end of its schedule. The run was stopped at step 215,000 of a planned 284,226. Its last evaluation is also its best, so WER had not started regressing when it stopped. The eight runs in the sweep stopped at different points (epoch 0.85 to 2.60 of their schedules), so the ranking below partly reflects how far each one got, not learning rate alone.
Compared with the other models
Ranked by WER this model is 3 of 8 in the sweep.
| model | params | WER | CER | latency | RTF |
|---|---|---|---|---|---|
| this model | 61.5 M | 8.34 | 3.01 | 1930 ms | 0.0732 |
| best moonshine-tiny fine-tune | 27.1 M | 9.95 | 3.91 | 942 ms | 0.0357 |
| whisper-large-v3-turbo + LoRA | 809 M | 4.12 | 1.78 | 7820 ms | 0.2968 |
All three decoded in the same harness, same clips, same CPU, batch 1. The fine-tuned Whisper is the more accurate model and will stay that way; what it costs is 4.1x the latency per utterance at 13x the parameters. Against the tiny fine-tune this model is 1.60 WER better for 2.0x the latency -- the base-vs-tiny trade this sweep exists to measure.
The untrained moonshine-base sits at 191.55 WER on this set: the stock
checkpoint is English-only, so it is a floor reference, not a baseline.
The sweep
| run | peak LR | warmup | epochs | checkpoint step | WER | CER |
|---|---|---|---|---|---|---|
| full_moonshine_base_all_data_wup1_ep3_lr5e-4 | 0.0005 | 1% | 3 | 244,000 / 284,226 | 7.36 | 2.87 |
| full_moonshine_base_all_data_wup1_ep3_lr1e-3 | 0.001 | 1% | 3 | 246,000 / 284,226 | 7.72 | 3.12 |
| full_moonshine_base_all_data_wup1_ep3_lr2e-4 (this model) | 0.0002 | 1% | 3 | 215,000 / 284,226 | 8.34 | 3.01 |
| full_moonshine_base_all_data_wup1_ep3_lr1e-4 | 0.0001 | 1% | 3 | 204,000 / 284,226 | 8.85 | 3.14 |
| full_moonshine_base_all_data_wup10_ep3_lr1e-4 | 0.0001 | 10% | 3 | 214,000 / 284,226 | 8.89 | 3.22 |
| full_moonshine_base_all_data_wup10_ep3_lr5e-5 | 5e-05 | 10% | 3 | 190,000 / 284,226 | 9.69 | 3.45 |
| full_moonshine_base_all_data_wup15_ep3_lr5e-5 | 5e-05 | 15% | 3 | 201,000 / 284,226 | 9.87 | 3.70 |
| full_moonshine_base_all_data_wup5_ep1_lr5e-5 | 5e-05 | 5% | 1 | 81,000 / 94,742 | 11.69 | 4.00 |
Training
| setting | value |
|---|---|
| learning rate | 0.0002 |
| warmup ratio | 0.01 |
| epochs | 3 |
| batch size | 32 x 1 accum |
| scheduler | cosine |
| weight decay | 0.01 |
| precision | bf16 |
| seed | 42 |
| total steps (planned) | 284,226 |
| steps completed | 215,000 |
Corpus (2.99M clips, each capped at 30 s): 2.96M clips, ~99% of
the mix), vezilka-asri (videa_so_transkript_od_yt (6,481), mozzila_common_voice (5,205),
alfa_audios (3,896), doniraj (3,157 accepted donations), sitel_audios
(3,136), fleurs_mk (1,853), jargon (1,274). The dialect repos and the held-out
test set are excluded by construction.
Usage
from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq
import librosa, torch
model_id = "h-gajdov/full_moonshine_base_all_data_wup1_ep3_lr2e-4"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForSpeechSeq2Seq.from_pretrained(model_id).eval()
wav, _ = librosa.load("clip.wav", sr=16000, mono=True)
feats = processor.feature_extractor([wav], sampling_rate=16000,
return_tensors="pt", padding=True)
with torch.no_grad():
ids = model.generate(**feats, max_length=192)
print(processor.batch_decode(ids, skip_special_tokens=True)[0])
max_length is capped at 192 because the decoder has 194 positions
(max_position_embeddings). Moonshine has no forced language/task tokens, so do
not pass language= or task= to generate() -- it will raise.
- Downloads last month
- 25
Model tree for h-gajdov/full_moonshine_base_all_data_wup1_ep3_lr2e-4
Base model
moonshine-ai/moonshine-base