Instructions to use mobilebytesensei/betterflow-en-streaming-fastconformer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use mobilebytesensei/betterflow-en-streaming-fastconformer with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("mobilebytesensei/betterflow-en-streaming-fastconformer") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
Betterflow — English streaming FastConformer, ONNX for sherpa-onnx
An ONNX export of nvidia/stt_en_fastconformer_hybrid_large_streaming_multi, prepared so it
loads in sherpa_onnx.OnlineRecognizer and produces live partials for English.
We are not the authors of the weights. Upstream is NVIDIA; this repo is a format conversion plus quantization.
Provenance and licence
| Upstream | nvidia/stt_en_fastconformer_hybrid_large_streaming_multi |
| Upstream licence | CC-BY-4.0 (read off the model card, not inferred) |
| This repo | CC-BY-4.0, inherited — attribution required |
| What changed | .nemo → ONNX via k2-fsa's own export path · set_default_att_context_size([70,13]) · int8 quantization |
| What did NOT change | the weights — no fine-tuning |
Please cite NVIDIA for the underlying model.
Contents — a transducer bundle (three graphs)
encoder.int8.onnx 131,507,640 B encoder.onnx 456,772,215 B
decoder.int8.onnx 3,955,863 B decoder.onnx 15,753,087 B
joiner.int8.onnx 1,408,183 B joiner.onnx 5,584,035 B
tokens.txt 11,896 B
int8 total ≈ 137 MB.
Measured
librispeech-en, n=50, through sherpa with a padded tail:
| WER | 7.7% pooled · 5.3% median |
| RTF | 0.021 |
| peak RSS | 662 MB |
| empty | 0/50 |
| script | 100% Latin |
Sample decode (int8, 2 s tail pad):
'concord returned to its place amidst the tents'
'congratulations were poured in upon the princess everywhere during her journey'
⚠️ Three things worth knowing
1. Pad the tail — and this bundle tells you exactly how much. sherpa's online recogniser only
decodes when num_frames_ready - num_processed >= window_size, and input_finished() does not
pad to a whole window, so up to window_size - 1 frames of every utterance are never decoded.
This encoder declares window_size = 121, chunk_shift = 112, subsampling_factor = 8 in its
ONNX metadata. At a 10 ms hop that is 1.21 s, so pad ≥ ~1.3 s; we use 2,000 ms. Shorter pads
lose words as deletions, which read as poor model quality rather than as a configuration error.
‼️ Read
window_sizeoff the graph rather than copying a number. An earlier version of this card quoted a0 → 20.7% · 500 → 10.4% · 2,000 → 5.4%sweep as if it were measured on this bundle. It was not — those are third-party figures from a streaming zipformer on Android, a different architecture whose chunk length we never read. The advice was right; the numbers were not ours.
2. downloadMb is not peakRssMb. 137 MB on disk, 662 MB resident — a 4.9× gap. Budget on
the resident figure.
3. Peak RSS is FLAT in utterance length — 671.2 MB at 5 s, 671.5 MB at 240 s. No utterance-length cap is needed for this bundle.
The reason is the cache-aware architecture — bounded left context plus a fixed cache — not the fact that it streams. ‼️ "Streaming ⇒ bounded memory" is false as a general rule: we measured a streaming decoder-only model whose peak RSS scales T^1.49 and walls at ~10.6 s of audio. Flat memory is a property of this family (cache-aware Conformer), and an offline Conformer's attention is O(T²). Check the scaling; do not infer it from the word "streaming."
Not evaluated
Device/Android verification · lookaheads other than [70,13] (0/80/480 ms are exportable via
the same script) · languages other than English · dictation-register audio — the numbers above are
read speech.