Surogate Speech | Training and Serving Engine | Toolkit and Evals |
License: CC-BY-NC-4.0 | Authors: Invergent

Jackrabbit 110M (Romanian)

Jackrabbit is Surogate's speech recognition family. This is the offline Romanian model: a 116M-parameter FastConformer with two decoders on one encoder, TDT and CTC, that writes cased, punctuated Romanian and runs comfortably on a CPU. For live audio use its sibling, Jackrabbit 110M Streaming.

It is served fastest by the surogate engine, which runs the encoder and both decoders natively, with no Python in the serving path. Every number on this page was measured, and the harnesses are named, so you can reproduce them.

Accuracy

Word error rate in %, lower is better. Every model ran through the same code on one RTX 5090 (bf16); other rows are our own runs of the public checkpoints, not their authors' figures.

FLEURS Romanian test (883 clips), with the Open ASR Leaderboard's own runner (nemo_asr/run_eval_ml.py, commit b2e6f04) and its multilingual normalizer:

Model Params WER Speed
Surogate Jackrabbit 110M, CTC + 4-gram 116M 5.69 2,531× real time
NVIDIA Canary 1B v2 1B 5.95 853×
Surogate Jackrabbit 110M, TDT greedy 116M 7.56 2,774×
OpenAI Whisper large-v3 1.55B 8.42 102×
SpeD 110M 110M 8.60 2,810×
NVIDIA Parakeet TDT 0.6B v3 0.6B 11.58 2,246×

The CTC + 4-gram row switches the decoder with a seven-line change to the runner, published with the results.

Common Voice 21 Romanian test (3,929 clips, the clip list from the SpeD paper), with the leaderboard normalizer and with SpeD's:

Model WER, leaderboard norm WER, SpeD norm
Surogate Jackrabbit 110M, CTC + 4-gram 2.07 2.19
Surogate Jackrabbit 110M, TDT greedy 2.19 2.30
SpeD 110M 3.47 3.58
NVIDIA Canary 1B v2 8.65 8.87
OpenAI Whisper large-v3 8.94 9.16
NVIDIA Parakeet TDT 0.6B v3 10.06 10.19

Our run of SpeD reproduces its published 3.57% on this set. Transcripts for every clip and model are in surogate-speech-evals.

Runs on

Runtime Where Decoding
surogate engine Linux, CPU or NVIDIA GPU TDT, CTC, CTC + 4-gram
NeMo (jackrabbit-110m-ro.nemo) Linux or macOS; CPU or GPU TDT, CTC, CTC + 4-gram
NeMo-Speech.cpp (gguf/) Linux CPU, no PyTorch TDT (7.99% on FLEURS under the same scoring as the NeMo file's 8.15%)

Running it

With surogate (recommended)

Needs surogate 1.5.4 or newer.

docker pull ghcr.io/invergent-ai/surogate:1.5.4
docker run --gpus all -p 8000:8000 ghcr.io/invergent-ai/surogate:1.5.4 \
  serve --stt surogate/jackrabbit-110m-ro --host 0.0.0.0 --port 8000

curl http://localhost:8000/v1/audio/transcriptions -F file=@interviu.wav -F model=surogate/jackrabbit-110m-ro

The API is OpenAI-compatible. Drop --gpus all to run on CPU. Server options are in the engine's speech docs.

With surogate-speech (local, CPU is fine)

pip install "surogate-speech[asr] @ git+https://github.com/invergent-ai/surogate-speech"
surogate-speech transcribe interviu.wav

With NeMo

from huggingface_hub import hf_hub_download
from nemo.collections.asr.models import ASRModel
from omegaconf import OmegaConf

model = ASRModel.restore_from(hf_hub_download("surogate/jackrabbit-110m-ro", "jackrabbit-110m-ro.nemo"))
model.change_decoding_strategy(decoder_type="rnnt")      # TDT greedy
print(model.transcribe(["audio.wav"])[0].text)

lm = hf_hub_download("surogate/jackrabbit-110m-ro", "lm-4gram-ro.nemo")   # loads in seconds
model.change_decoding_strategy(OmegaConf.create({         # CTC + 4-gram, the best accuracy
    "strategy": "beam_batch",
    "beam": {"beam_size": 32, "ngram_lm_model": lm, "ngram_lm_alpha": 0.5, "beam_beta": 2.0,
             "return_best_hypothesis": True, "allow_cuda_graphs": False},
}), decoder_type="ctc")
print(model.transcribe(["audio.wav"])[0].text)

Files

file size what it is
jackrabbit-110m-ro.nemo 466 MB the model: FastConformer encoder (17 layers, width 512, 80 ms frames), TDT and CTC decoders, 2,048-piece SentencePiece vocabulary with case and punctuation
lm-4gram-ro.nemo 2.07 GB the 4-gram language model for CTC beam search, in NeMo's fast-loading format
lm-4gram-ro.arpa 2.28 GB the same language model as standard ARPA, for other decoders
gguf/ 257 MB (F16), 150 MB (Q8_0) the model for NeMo-Speech.cpp, TDT decoder

The language model is built over the model's own 2,048 SentencePiece pieces, so it cannot be paired with another tokenizer. Input is 16 kHz mono; the engine and surogate-speech resample other formats.

Limitations

Romanian only. Accuracy was measured on read and parliamentary speech; telephone audio, heavy noise, overlapping speakers and strong regional accents are not measured here. Numbers, acronyms and rare proper names cause a large share of the remaining errors. The model transcribes; it does not identify speakers.

License

CC-BY-NC-4.0: free for research and other non-commercial use; for commercial use, contact Invergent. Fine-tuned from nvidia/parakeet-tdt_ctc-110m by NVIDIA (CC-BY-4.0).

Downloads last month
51
GGUF
Model size
0.1B params
Architecture
asr
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for surogate/jackrabbit-110m-ro

Quantized
(23)
this model

Collection including surogate/jackrabbit-110m-ro

Paper for surogate/jackrabbit-110m-ro

Evaluation results