Phi-4-multimodal-instruct, audio-to-text carve-out, ONNX

This repository holds only the speech transcription path of microsoft/Phi-4-multimodal-instruct, exported to ONNX for onnx-asr.

The source model is a full multimodal LLM: a conformer audio encoder and a SigLIP vision tower share one Phi-4-mini language model, which is specialised at run time by a vision LoRA adapter or a speech LoRA adapter. This export takes the audio branch only:

  • the speech LoRA adapter is merged into the language model, proven character-identical to the adapter path on four FLEURS clips before export;
  • the vision tower, the vision LoRA adapter and the vision audio projector are dropped entirely.

This is an audio-only carve-out. It transcribes speech. It cannot see images, answer questions, translate or hold a conversation.

Original model by Microsoft, MIT licence. This export keeps that licence. See REPORT.md for the full export record, the merge proof, the numerical error bounds and the deviations.

Usage

The model loads on the feat/speech-llm-qwen3-asr branch of the TigreGotico onnx-asr fork. No runtime changes were needed: the export is entirely config-driven.

pip install "onnx-asr @ git+https://github.com/TigreGotico/onnx-asr@feat/speech-llm-qwen3-asr"
import onnx_asr

model = onnx_asr.load_model("OpenVoiceOS/phi-4-multimodal-asr-onnx")
print(model.recognize("clip.wav"))

# roughly 3x faster and 4x smaller
model = onnx_asr.load_model("OpenVoiceOS/phi-4-multimodal-asr-onnx", quantization="int8")

Languages

Phi-4-multimodal supports eight languages for audio: English, Chinese, German, French, Italian, Japanese, Spanish and Portuguese. Other languages are out of domain.

Graphs

graph inputs outputs
encoder.onnx input_features (1, N) f32, raw 16 kHz waveform audio_embeds (1, L, 3072) f32
embed_tokens.onnx input_ids (1, S) i64 inputs_embeds (1, S, 3072) f32
decoder.onnx inputs_embeds, attn_bias, position_ids, 32 KV pairs logits, 32 KV pairs

The Phi-4-multimodal feature extractor is the nonstandard SpeechLib log-filterbank, so it is baked into the encoder graph and the model declares "preprocessor": "identity". The runtime passes the raw waveform straight through.

Two things to know

Roughly 40 s of audio per call. The upstream conformer switches to a chunked attention path above 500 post-CNN frames, and that branch cannot be expressed in a single ONNX graph. The exported encoder reproduces the unchunked path exactly, which is upstream's own stated design range. Use VAD segmentation for longer audio.

int8 excludes the encoder. Quantizing the conformer breaks transcription badly, so encoder_int8.onnx is a full-precision copy of encoder.onnx. Only embed_tokens and decoder are actually quantized. The graph-swap evidence is in REPORT.md.

Parity

ONNX fp32 is character-identical to native transformers on all four FLEURS clips (two English, two Portuguese). The int8 variant is faithful, with small punctuation and article differences on three of the four clips.

Downloads last month
104
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OpenVoiceOS/phi-4-multimodal-asr-onnx

Quantized
(9)
this model

Collections including OpenVoiceOS/phi-4-multimodal-asr-onnx