typhoon-whisper-turbo, ONNX with word timing

ONNX conversion of typhoon-ai/typhoon-whisper-turbo (MIT) for running in the browser with transformers.js. The decoder also outputs its cross-attention maps, which word-level timing is computed from.

Changes from the original files:

  • generation_config.json: alignment_heads set to [[2,11],[1,1],[3,6],[2,4],[3,14],[2,8]], the heads that track time best after fine-tuning.
  • Weights converted to ONNX in 4-bit and 8-bit variants. No retraining.

The model was trained without timestamp tokens: generate with return_timestamps: false and return_token_timestamps: true.

All credit for the model goes to its original authors (typhoon-ai).

Downloads last month
87
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 2nuttertools/typhoon-whisper-turbo

Quantized
(4)
this model