Whisper-tiny INT8 ONNX
Dynamic INT8 quantized weights for openai/whisper-tiny, optimized for low-latency, edge-native speech recognition directly in the browser via ONNX Runtime WebAssembly.
Model Summary
- Base Model:
openai/whisper-tiny - Quantization Method: Dynamic INT8 (
QuantType.QUInt8) - Total Repository Size: ~103.4 MB
- Active Runtime Footprint: ~56 MB (Only loads Encoder + KV-cache Decoder during streaming inference)
- Intended Use: Client-side ASR execution for
edgespeech-wasm
Files & Sizes
encoder_model_int8.onnx(~9.6 MB): Audio feature extractor & encoderdecoder_model_int8.onnx(~47.5 MB): Standard autoregressive text decoderdecoder_with_past_model_int8.onnx(~46.3 MB): Decoder optimized with KV-cache for low-latency streaming
- Downloads last month
- -
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support