Moonshine Streaming tiny-es (ONNX, incremental five-graph runtime)

ONNX export of moonshine-ai/moonshine-streaming-tiny-es (Moonshine v2, "Ergodic Streaming Encoder"), split into the five graphs the official Moonshine streaming runtime drives, so audio is encoded incrementally as it arrives instead of re-encoding the whole utterance:

Graph Inputs -> outputs
frontend.onnx audio_chunk[1,N] (N = multiple of 640 samples) + 5 carried states -> features[1,N/320,320] + updated states
encoder.onnx / encoder_int8.onnx features[1,T,320] -> encoded[1,T,320] (sliding-window attention; run on a window of total_left_context=96 past frames + new frames, keep frames older than total_lookahead=16 as stable)
adapter.onnx / adapter_int8.onnx encoded, pos_offset -> memory[1,T,320] (adds absolute position embeddings, 4096 positions = 82 s per segment)
cross_kv.onnx / cross_kv_int8.onnx memory -> k_cross, v_cross [6,1,8,M,40]
decoder_kv.onnx / decoder_kv_int8.onnx token[1,L], self K/V cache (length may be 0), cross K/V -> logits, updated self K/V

streaming_config.json carries the dimensions, BOS/EOS ids, frontend state shapes, the per-layer attention windows and the derived total_lookahead / total_left_context.

Export notes:

  • Exported with the official moonshine/scripts/export.py recipe (torch.onnx, opset 17).
  • Sliding windows use the inclusive semantics the models were trained with (and that the official runtime uses); for the multilingual checkpoints the HF config's (17, 5)/(17, 1) notation is normalized to the inclusive (16, 4)/(16, 0). The fp32 graphs reproduce transformers' MoonshineStreamingForConditionalGeneration greedy output token-for-token when its mask uses the same inclusive windows.
  • *_int8.onnx: onnxruntime.quantization.quantize_dynamic, QInt8 per-channel weights on MatMul/Gemm with constant B. The frontend is shipped fp32 only (it is ~3% of compute and its convolution state carry should stay exact).

Verification

Greedy decoding, CPU, on 20 FLEURS es_419 test. The fp32 five-graph pipeline is token-for-token identical to transformers on every utterance of the set.

Runtime WER fp32 WER int8
onnxruntime (Python, whole utterance) 6.68% 6.47%
WinSTT Rust engine (incremental, energy endpointer) 6.89% -

Rust fp32 real-time factor on a desktop CPU (under concurrent load): 0.039. The set is small (20-41 utterances), so treat the numbers as a regression check, not a benchmark. The encoder graph does not run on DirectML (ORT 1.24 DML EP rejects its attention-head Reshape), so use the CPU execution provider.

License

MIT, inherited from moonshine-ai/moonshine-streaming-tiny-es. See LICENSE. Tokenizer and configuration files are copied unchanged from the base repository.

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Masterx/moonshine-streaming-tiny-es-ONNX

Quantized
(2)
this model