Moonshine Streaming tiny-es (ONNX, incremental five-graph runtime)
ONNX export of moonshine-ai/moonshine-streaming-tiny-es (Moonshine v2, "Ergodic Streaming Encoder"), split into the five graphs the official Moonshine streaming runtime drives, so audio is encoded incrementally as it arrives instead of re-encoding the whole utterance:
| Graph | Inputs -> outputs |
|---|---|
frontend.onnx |
audio_chunk[1,N] (N = multiple of 640 samples) + 5 carried states -> features[1,N/320,320] + updated states |
encoder.onnx / encoder_int8.onnx |
features[1,T,320] -> encoded[1,T,320] (sliding-window attention; run on a window of total_left_context=96 past frames + new frames, keep frames older than total_lookahead=16 as stable) |
adapter.onnx / adapter_int8.onnx |
encoded, pos_offset -> memory[1,T,320] (adds absolute position embeddings, 4096 positions = 82 s per segment) |
cross_kv.onnx / cross_kv_int8.onnx |
memory -> k_cross, v_cross [6,1,8,M,40] |
decoder_kv.onnx / decoder_kv_int8.onnx |
token[1,L], self K/V cache (length may be 0), cross K/V -> logits, updated self K/V |
streaming_config.json carries the dimensions, BOS/EOS ids, frontend state shapes, the
per-layer attention windows and the derived total_lookahead / total_left_context.
Export notes:
- Exported with the official
moonshine/scripts/export.pyrecipe (torch.onnx, opset 17). - Sliding windows use the inclusive semantics the models were trained with (and that the
official runtime uses); for the multilingual checkpoints the HF config's
(17, 5)/(17, 1)notation is normalized to the inclusive(16, 4)/(16, 0). The fp32 graphs reproducetransformers'MoonshineStreamingForConditionalGenerationgreedy output token-for-token when its mask uses the same inclusive windows. *_int8.onnx:onnxruntime.quantization.quantize_dynamic, QInt8 per-channel weights on MatMul/Gemm with constant B. The frontend is shipped fp32 only (it is ~3% of compute and its convolution state carry should stay exact).
Verification
Greedy decoding, CPU, on 20 FLEURS es_419 test. The fp32 five-graph pipeline is token-for-token identical to transformers on every utterance of the set.
| Runtime | WER fp32 | WER int8 |
|---|---|---|
| onnxruntime (Python, whole utterance) | 6.68% | 6.47% |
| WinSTT Rust engine (incremental, energy endpointer) | 6.89% | - |
Rust fp32 real-time factor on a desktop CPU (under concurrent load): 0.039. The set is small (20-41 utterances), so treat the numbers as a regression check, not a benchmark. The encoder graph does not run on DirectML (ORT 1.24 DML EP rejects its attention-head Reshape), so use the CPU execution provider.
License
MIT, inherited from moonshine-ai/moonshine-streaming-tiny-es. See LICENSE.
Tokenizer and configuration files are copied unchanged from the base repository.
- Downloads last month
- 10
Model tree for Masterx/moonshine-streaming-tiny-es-ONNX
Base model
moonshine-ai/moonshine-streaming-tiny-es