Clipsun AI model files
Model files downloaded by the Clipsun AI desktop app on first run. They are format conversions of biodatlab/whisper-th-large-v3-combined (Thonburian Whisper, Thai fine-tune of OpenAI Whisper large-v3), licensed Apache-2.0. The weights are unchanged apart from the format and float16 precision.
| Folder | Format | Used on |
|---|---|---|
whisper-th-large-v3-combined/ |
CTranslate2, float16 (for faster-whisper) | NVIDIA graphics cards (CUDA) |
whisper-th-large-v3-combined-onnx/ |
ONNX, float16, fixed-shape encoder + decoder graphs for ONNX Runtime DirectML | AMD / Intel graphics cards |
The ONNX decoder takes one token per beam for one 15-second audio chunk (5 beams, 256
token slots) with per-layer key/value caches, and also returns the cross-attention of the
alignment heads (align_heads.json) for word timestamps.
Credits
- Thonburian Whisper by Biomedical and Data Lab, Mahidol University (biodatlab), Apache-2.0.
- OpenAI Whisper, MIT.
See LICENSE (Apache License 2.0).
Model tree for qroudry/clipsun-models
Base model
openai/whisper-large-v3 Finetuned
biodatlab/whisper-th-large-v3-combined