KOTube models
Models used by the KOTube browser extension (karaoke on YouTube), converted for onnxruntime-web (WebGPU). The extension downloads them once and keeps them on the user's disk.
KOTubeSub 2 β Thai lyrics recognition (current)
A format conversion of typhoon-ai/typhoon-whisper-turbo (revision 3c03fa8), stored in FP16 and split for step-by-step greedy decoding in the browser. No new training; all credit for the model goes to its authors. License: MIT (OpenTyphoon terms also apply to the original model).
kotubesub2/encoder-fp16.onnx:mel [1, 128, 3000](float32, a 30 s Whisper log-mel window) βcross_k,cross_v [4, 1, 20, 1500, 64](float16): the encoder followed by each decoder layer's cross-attention key/value projections.kotubesub2/decoder-fp16.onnx: one decoding step fornnew tokens:tokens [1, n]int64,offset [1]int64,past_k,past_v [4, 1, 20, P, 64]float16,mask [1, 1, n, P + n]float32,cross_k,cross_vβlogits [1, 51866],present_k,present_v,align [6, n, 1500](cross-attention of the alignment heads, for token times by DTW).kotubesub2/tokens.json: token id β its bytes (as a latin1 string).kotubesub/mel-filters.f32is shared with KOTubeSub 1 (Whisper's 128-bin filterbank).
Export and checks: research/live-lyrics/export_turbo.py and check_turbo.py in the KOTube repository.
KOTubeSub English β English lyrics recognition (optional)
The same format as KOTubeSub 2, converted from openai/whisper-large-v3-turbo
(MIT) and decoded with the English language token; no new training. Used for videos the user marks as English.
Shares kotubesub2/tokens.json and kotubesub/mel-filters.f32 with KOTubeSub 2.
kotubesub-en/encoder-fp16.onnx,kotubesub-en/decoder-fp16.onnx: inputs and outputs as KOTubeSub 2
KOTubeSub 1 β Thai lyrics recognition (older extension versions)
kotubesub/kotubesub-fp16.onnx: input a 30 s Whisper log-mel window mel [1, 128, 3000] (float32), output CTC
logits logits [1, 1500, 80] (float32, 20 ms per frame; greedy decode with vocab.json, blank = 0).
mel-filters.f32 is Whisper's 128-bin filterbank (201 x 128, float32, row-major).
KOTubeSub is a format conversion, not new training: the encoder of typhoon-ai/typhoon-whisper-large-v3 (revision 748e8a4) and the CTC head of typhoon-ai/typhoon-whisper-large-v3-ctc, merged into one graph and stored in FP16. All credit for the models goes to their authors.
- Encoder: MIT (OpenTyphoon terms also apply to the original model)
- CTC head and vocabulary: Apache-2.0
Vocal removal
separation/UVR-MDX-NET-Inst_HQ_5.onnx: UVR MDX-Net Inst HQ 5 from
TRvlvr/model_repo by Anjok07 and the
UVR team (MIT). Weights unchanged; only the time axis of the
input/output was made symbolic.