Instructions to use surogate/jackrabbit-110m-ro with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use surogate/jackrabbit-110m-ro with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("surogate/jackrabbit-110m-ro") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
Surogate Speech |
Training and Serving Engine |
Toolkit and Evals |
License: CC-BY-NC-4.0 | Authors: Invergent
Jackrabbit 110M (Romanian)
Jackrabbit is Surogate's speech recognition family. This is the offline Romanian model: a 116M-parameter FastConformer with two decoders on one encoder, TDT and CTC, that writes cased, punctuated Romanian and runs comfortably on a CPU. For live audio use its sibling, Jackrabbit 110M Streaming.
It is served fastest by the surogate engine, which runs the encoder and both decoders natively, with no Python in the serving path. Every number on this page was measured, and the harnesses are named, so you can reproduce them.
Accuracy
Word error rate in %, lower is better. Every model ran through the same code on one RTX 5090 (bf16); other rows are our own runs of the public checkpoints, not their authors' figures.
FLEURS Romanian test (883 clips), with the Open ASR Leaderboard's own runner
(nemo_asr/run_eval_ml.py, commit b2e6f04) and its multilingual normalizer:
| Model | Params | WER | Speed |
|---|---|---|---|
| Surogate Jackrabbit 110M, CTC + 4-gram | 116M | 5.69 | 2,531× real time |
| NVIDIA Canary 1B v2 | 1B | 5.95 | 853× |
| Surogate Jackrabbit 110M, TDT greedy | 116M | 7.56 | 2,774× |
| OpenAI Whisper large-v3 | 1.55B | 8.42 | 102× |
| SpeD 110M | 110M | 8.60 | 2,810× |
| NVIDIA Parakeet TDT 0.6B v3 | 0.6B | 11.58 | 2,246× |
The CTC + 4-gram row switches the decoder with a seven-line change to the runner, published with the results.
Common Voice 21 Romanian test (3,929 clips, the clip list from the SpeD paper), with the leaderboard normalizer and with SpeD's:
| Model | WER, leaderboard norm | WER, SpeD norm |
|---|---|---|
| Surogate Jackrabbit 110M, CTC + 4-gram | 2.07 | 2.19 |
| Surogate Jackrabbit 110M, TDT greedy | 2.19 | 2.30 |
| SpeD 110M | 3.47 | 3.58 |
| NVIDIA Canary 1B v2 | 8.65 | 8.87 |
| OpenAI Whisper large-v3 | 8.94 | 9.16 |
| NVIDIA Parakeet TDT 0.6B v3 | 10.06 | 10.19 |
Our run of SpeD reproduces its published 3.57% on this set. Transcripts for every clip and model are in surogate-speech-evals.
Runs on
| Runtime | Where | Decoding |
|---|---|---|
| surogate engine | Linux, CPU or NVIDIA GPU | TDT, CTC, CTC + 4-gram |
NeMo (jackrabbit-110m-ro.nemo) |
Linux or macOS; CPU or GPU | TDT, CTC, CTC + 4-gram |
NeMo-Speech.cpp (gguf/) |
Linux CPU, no PyTorch | TDT (7.99% on FLEURS under the same scoring as the NeMo file's 8.15%) |
Running it
With surogate (recommended)
Needs surogate 1.5.4 or newer.
docker pull ghcr.io/invergent-ai/surogate:1.5.4
docker run --gpus all -p 8000:8000 ghcr.io/invergent-ai/surogate:1.5.4 \
serve --stt surogate/jackrabbit-110m-ro --host 0.0.0.0 --port 8000
curl http://localhost:8000/v1/audio/transcriptions -F file=@interviu.wav -F model=surogate/jackrabbit-110m-ro
The API is OpenAI-compatible. Drop --gpus all to run on CPU. Server options are in the
engine's speech docs.
With surogate-speech (local, CPU is fine)
pip install "surogate-speech[asr] @ git+https://github.com/invergent-ai/surogate-speech"
surogate-speech transcribe interviu.wav
With NeMo
from huggingface_hub import hf_hub_download
from nemo.collections.asr.models import ASRModel
from omegaconf import OmegaConf
model = ASRModel.restore_from(hf_hub_download("surogate/jackrabbit-110m-ro", "jackrabbit-110m-ro.nemo"))
model.change_decoding_strategy(decoder_type="rnnt") # TDT greedy
print(model.transcribe(["audio.wav"])[0].text)
lm = hf_hub_download("surogate/jackrabbit-110m-ro", "lm-4gram-ro.nemo") # loads in seconds
model.change_decoding_strategy(OmegaConf.create({ # CTC + 4-gram, the best accuracy
"strategy": "beam_batch",
"beam": {"beam_size": 32, "ngram_lm_model": lm, "ngram_lm_alpha": 0.5, "beam_beta": 2.0,
"return_best_hypothesis": True, "allow_cuda_graphs": False},
}), decoder_type="ctc")
print(model.transcribe(["audio.wav"])[0].text)
Files
| file | size | what it is |
|---|---|---|
jackrabbit-110m-ro.nemo |
466 MB | the model: FastConformer encoder (17 layers, width 512, 80 ms frames), TDT and CTC decoders, 2,048-piece SentencePiece vocabulary with case and punctuation |
lm-4gram-ro.nemo |
2.07 GB | the 4-gram language model for CTC beam search, in NeMo's fast-loading format |
lm-4gram-ro.arpa |
2.28 GB | the same language model as standard ARPA, for other decoders |
gguf/ |
257 MB (F16), 150 MB (Q8_0) | the model for NeMo-Speech.cpp, TDT decoder |
The language model is built over the model's own 2,048 SentencePiece pieces, so it cannot be paired with another tokenizer. Input is 16 kHz mono; the engine and surogate-speech resample other formats.
Limitations
Romanian only. Accuracy was measured on read and parliamentary speech; telephone audio, heavy noise, overlapping speakers and strong regional accents are not measured here. Numbers, acronyms and rare proper names cause a large share of the remaining errors. The model transcribes; it does not identify speakers.
License
CC-BY-NC-4.0: free for research and other non-commercial use; for commercial use, contact Invergent. Fine-tuned from nvidia/parakeet-tdt_ctc-110m by NVIDIA (CC-BY-4.0).
- Downloads last month
- 51
8-bit
16-bit
Model tree for surogate/jackrabbit-110m-ro
Base model
nvidia/parakeet-tdt_ctc-110mCollection including surogate/jackrabbit-110m-ro
Paper for surogate/jackrabbit-110m-ro
Evaluation results
- WER (CTC + 4-gram) on FLEURS ro_ro test (Open ASR Leaderboard protocol)test set self-reported5.690
- WER (TDT greedy) on FLEURS ro_ro test (Open ASR Leaderboard protocol)test set self-reported7.560
- WER (CTC + 4-gram) on Common Voice 21 ro test (SpeD clip list)test set self-reported2.070
- WER (TDT greedy) on Common Voice 21 ro test (SpeD clip list)test set self-reported2.190