Whisper-tiny INT8 ONNX

Dynamic INT8 quantized weights for openai/whisper-tiny, optimized for low-latency, edge-native speech recognition directly in the browser via ONNX Runtime WebAssembly.

Model Summary

  • Base Model: openai/whisper-tiny
  • Quantization Method: Dynamic INT8 (QuantType.QUInt8)
  • Total Repository Size: ~103.4 MB
  • Active Runtime Footprint: ~56 MB (Only loads Encoder + KV-cache Decoder during streaming inference)
  • Intended Use: Client-side ASR execution for edgespeech-wasm

Files & Sizes

  • encoder_model_int8.onnx (~9.6 MB): Audio feature extractor & encoder
  • decoder_model_int8.onnx (~47.5 MB): Standard autoregressive text decoder
  • decoder_with_past_model_int8.onnx (~46.3 MB): Decoder optimized with KV-cache for low-latency streaming
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support