distill-whisper-th-small, ONNX with word timing

ONNX conversion of biodatlab/distill-whisper-th-small (MIT) for running in the browser with transformers.js. The decoder also outputs its cross-attention maps, which word-level timing is computed from.

Changes from the original files:

  • generation_config.json: alignment_heads set to [[2,4],[2,8],[1,10]]. The original lists heads of the 12-layer whisper-small decoder; this distilled model has 4 decoder layers.
  • Weights converted to ONNX in half precision, 4-bit and 8-bit variants. No retraining.

The model was trained without timestamp tokens: generate with return_timestamps: false and return_token_timestamps: true.

All credit for the model goes to its original authors (biodatlab).

Downloads last month
48
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 2nuttertools/distill-whisper-th-small

Quantized
(2)
this model