granite-speech-4.1-2b-nar-onnx

ONNX export of ibm-granite/granite-speech-4.1-2b-nar for onnx-asr.

This model does not decode token by token. A conformer encoder with a BPE CTC head writes a first-pass hypothesis, and a bidirectional Granite-4.0-1b editor rewrites that hypothesis in one forward pass. There is no KV cache and no decoding loop, which makes it the fastest fp32 CPU model in this collection: RTFx 2.9 - 3.9 on a 6-core desktop CPU, against 0.28 - 1.3 for the autoregressive exports.

Languages: English, French, German, Spanish, Portuguese.

Requires the feat/granite-nar branch

The granite-nar model type is not in upstream onnx-asr yet. Install the fork:

pip install "onnx-asr @ git+https://github.com/TigreGotico/onnx-asr@feat/granite-nar"

Usage

import onnx_asr

model = onnx_asr.load_model("granite-nar", "OpenVoiceOS/granite-speech-4.1-2b-nar-onnx")
print(model.recognize("audio.wav"))

For the smaller int8 graphs:

model = onnx_asr.load_model(
    "granite-nar", "OpenVoiceOS/granite-speech-4.1-2b-nar-onnx", quantization="int8"
)

Which files to use

Configuration Size Quality on the test clips
fp32 (default) 9.8 GB character-identical to the PyTorch model on all 4 clips
int8 encoder + fp32 editor 3.1 GB both English clips exact, two Portuguese words softened
int8 (both graphs) 2.5 GB 1 of 4 clips exact, the rest lose or garble words

fp32 is also the fastest configuration. Unlike the autoregressive models in this collection, int8 does not speed this one up: a single editor pass is dominated by large GEMMs, where the dequantize overhead is not repaid.

The int8 loss is almost entirely in the editor, which had to be quantized per tensor rather than per channel (per-channel quantize_dynamic needed more than 110 GB of RAM on this graph). The encoder is per-channel with its in-graph feature extraction kept in fp32, and reaches cosine 0.98 against fp32.

Graphs

File Inputs Outputs
encoder.onnx input_features (1, N) raw 16 kHz waveform audio_embeds (1, L, 2048), ctc_logits (1, C, 100352), audio_embeds_lens (1,), ctc_lens (1,)
editor.onnx audio_embeds (1, L, 2048), text_ids (1, T) logits (1, T, 100352)

The encoder takes the raw waveform: the torchaudio log-mel front end is computed inside the graph as a strided convolution against a windowed DFT basis, to a maximum absolute error of 8.3e-7 against the transformers feature extractor. The editor holds the token embedding table, which is tied to its output head, so its text input is token ids rather than embeddings.

Between the graphs the runtime does the CTC collapse, opens a blank editing slot around every surviving token, and collapses again after the editor.

Accuracy

FLEURS validation, 2 English and 2 Portuguese clips, against the PyTorch model (trust_remote_code=True, fp32, CPU):

Clip ONNX fp32
en_1 identical
en_2 identical
pt_1 identical
pt_2 identical

The full transcripts, timings, the int8 comparison and the graph-level checks are in REPORT.md.

Limitations

  • Batch 1. recognize_batch loops over the waveforms.
  • Greedy CTC decoding, no timestamps, no beam search.
  • Transcription only, and no language selection.

Licence

Apache-2.0, same as the source model.

Downloads last month
1,249
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OpenVoiceOS/granite-speech-4.1-2b-nar-onnx

Quantized
(10)
this model

Collections including OpenVoiceOS/granite-speech-4.1-2b-nar-onnx