Phi-4-multimodal-instruct, audio-to-text carve-out, ONNX
This repository holds only the speech transcription path of
microsoft/Phi-4-multimodal-instruct,
exported to ONNX for onnx-asr.
The source model is a full multimodal LLM: a conformer audio encoder and a SigLIP vision tower share one Phi-4-mini language model, which is specialised at run time by a vision LoRA adapter or a speech LoRA adapter. This export takes the audio branch only:
- the speech LoRA adapter is merged into the language model, proven character-identical to the adapter path on four FLEURS clips before export;
- the vision tower, the vision LoRA adapter and the vision audio projector are dropped entirely.
This is an audio-only carve-out. It transcribes speech. It cannot see images, answer questions, translate or hold a conversation.
Original model by Microsoft, MIT licence. This export keeps that licence.
See REPORT.md for the full export record, the merge proof, the
numerical error bounds and the deviations.
Usage
The model loads on the feat/speech-llm-qwen3-asr branch of the TigreGotico
onnx-asr fork. No runtime changes were needed: the export is entirely
config-driven.
pip install "onnx-asr @ git+https://github.com/TigreGotico/onnx-asr@feat/speech-llm-qwen3-asr"
import onnx_asr
model = onnx_asr.load_model("OpenVoiceOS/phi-4-multimodal-asr-onnx")
print(model.recognize("clip.wav"))
# roughly 3x faster and 4x smaller
model = onnx_asr.load_model("OpenVoiceOS/phi-4-multimodal-asr-onnx", quantization="int8")
Languages
Phi-4-multimodal supports eight languages for audio: English, Chinese, German, French, Italian, Japanese, Spanish and Portuguese. Other languages are out of domain.
Graphs
| graph | inputs | outputs |
|---|---|---|
encoder.onnx |
input_features (1, N) f32, raw 16 kHz waveform |
audio_embeds (1, L, 3072) f32 |
embed_tokens.onnx |
input_ids (1, S) i64 |
inputs_embeds (1, S, 3072) f32 |
decoder.onnx |
inputs_embeds, attn_bias, position_ids, 32 KV pairs |
logits, 32 KV pairs |
The Phi-4-multimodal feature extractor is the nonstandard SpeechLib
log-filterbank, so it is baked into the encoder graph and the model declares
"preprocessor": "identity". The runtime passes the raw waveform straight
through.
Two things to know
Roughly 40 s of audio per call. The upstream conformer switches to a chunked attention path above 500 post-CNN frames, and that branch cannot be expressed in a single ONNX graph. The exported encoder reproduces the unchunked path exactly, which is upstream's own stated design range. Use VAD segmentation for longer audio.
int8 excludes the encoder. Quantizing the conformer breaks transcription
badly, so encoder_int8.onnx is a full-precision copy of encoder.onnx. Only
embed_tokens and decoder are actually quantized. The graph-swap evidence is
in REPORT.md.
Parity
ONNX fp32 is character-identical to native transformers on all four FLEURS clips (two English, two Portuguese). The int8 variant is faithful, with small punctuation and article differences on three of the four clips.
- Downloads last month
- 104
Model tree for OpenVoiceOS/phi-4-multimodal-asr-onnx
Base model
microsoft/Phi-4-multimodal-instruct