sl-caribou
Whisper-small continued on the 2,147.6h khazina-mix (real Tajik speech; 860h of it carries Whisper-generated pseudo-labels). 2 full epochs, 33856/33856 steps. FLEURS-tg n=600 beam5 judge-norm WER 11.89 / CER 4.41, vs its own baseline 13.17 (-1.28 pp). IMPORTANT CAVEAT: this was the in-family probe of a self-training experiment -- the labels came from a Whisper champion, and the cross-family probe (Parakeet, same data) did NOT improve. Part of that gain may be the model agreeing with its own labeller's errors rather than learning new signal. Treat as a distillation artifact, not as evidence the data is good. See the run's NEEDS_TOHIR.md.
- Downloads last month
- -
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support