TIGER for Divide and Remaster, ONNX (int8 weights)
TIGER (Xu, Li, Chen, Hu, "TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech
Separation", ICLR 2025, arXiv:2410.01469), its Divide and Remaster model TIGERDNR:
three band-split models of 1.4 M parameters (57 bands, multi-scale convolutions, frame and frequency attention), each
separating three sources and keeping one. It splits a soundtrack into three stems: dialogue, music, effects. This is the
three networks between their STFT and iSTFT as one graph, as
@audio/neural-separate runs it (model: 'tiger'),
converted from JusperLee/TIGER-DnR.
| File | Size | SHA-256 |
|---|---|---|
tiger.int8.onnx |
7.3 MB (7,320,735 bytes) | 2189859ac7e89cf4d9d284c6f393e737aaa2b3233ed7ee424ad622012bce7bca |
The float32 export it is made from is 28.6 MB, of which 11.7 MB were node names; float16 weights would be 10.9 MB.
Source
- Code: JusperLee/TIGER at
9f18d4a10a7137e1ce8052cfb62215179f1287b6:look2hear/models/tiger_dnr.pyand the layer modules it reads. - Weights: JusperLee/TIGER-DnR at revision
b7a59560bbca10febbcd46fb01600f868e587f57,model.safetensors(SHA-256dd1c696e72f6adea0085ef1af640882a8260519ad666422835e387a5b4abdd2a).
Licence and attribution
Apache License 2.0 (LICENSE): the weights' model card states license: apache-2.0 (the code repository's
LICENSE is MIT, its README badge Apache 2.0). Modified from the original weights: converted to ONNX with the three
models in one graph; F.adaptive_avg_pool1d exported as a cumulative sum read at torch's windows, UConvBlock's global
sum started at its first term and each GroupNorm's statistics reduced axis by axis (none changing the function); the
weights stored in 8 bits (below).
TIGER by Mohan Xu, Kai Li, Guo Chen and Xiaolin Hu (Tsinghua University); ONNX export and compaction by audiojs. Trained on Divide and Remaster v1, sourced as v2 is (FMA and FSD50K clips, each under its own licence).
@inproceedings{xu2025tiger,
title = {TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation},
author = {Xu, Mohan and Li, Kai and Chen, Guo and Hu, Xiaolin},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2025},
eprint = {2410.01469},
archivePrefix = {arXiv}
}
Graph
One mono 12 s segment (529,200 samples at 44.1 kHz, 1034 frames) per run, the three models in one graph:
| name | shape | ||
|---|---|---|---|
| input | mix_spec |
[1, 2, 1025, 1034] | STFT: n 2048, hop 512, Hann, unnormalized, centered with reflect padding; re, im |
| output | stems_spec |
[1, 6, 1025, 1034] | dialogue, music, effects, each re and im |
Segments start every 4 s, unweighted, zero-padded past the ends, their sum divided by three
(TIGERDNR.wav_chunk_inference).
import separate from '@audio/neural-separate'
let { stems } = await separate([mono], { sampleRate: 48000, model: 'tiger' }) // stems.dialogue, .music, .effects
Export, compaction, verification
scripts/export-tiger.py --verify exports the graph and compares it with the three TIGER.forward on one segment of
tones and noise: max |diff| ≤ 4.2e-7 of max |y|; the package's pipeline matches TIGERDNR.wav_chunk_inference to
95–140 dB SNR per stem. scripts/compact.py --model tiger --calibrate <a Divide and Remaster v3 tuning clip> makes this
file from it:
- Weights: all 444 convolution weights of over 1024 values stored in int8 (symmetric, a scale per output channel): rounded together they move the output 41.4 dB under its power on the calibration clip, within the 40 dB the script allows before keeping any in float16.
- Computed in float32 on every backend: each weight is Cast and multiplied by its scale in the graph, which onnxruntime folds at load (the session holds float32 weights).
- Folded and named short, changing no value: with the input's shape fixed, every value computable from the weights and the shapes alone is stored as the graph computes it (56,652 nodes to 38,817); node and value names are base-36 counters (the graph without its weights 11.7 MB to 1.9 MB).
- Against the export on the calibration clip: max |diff| 1.7e-2 of max |y|, SNR 41.4 dB.
Quality
The 30 test clips @audio/neural-separate measures TIGER on (Divide and Remaster v3's English test set, every 40th of its 1200 clips of 60 s; Watcharasupat, Wu, Orife, 2024, CC BY-SA 4.0): each stem against its reference over the whole clip, SNR, the median over clips, dB:
| dialogue | music | effects | |
|---|---|---|---|
| export (float32) | 12.68 | 10.24 | 8.06 |
| this file | 12.66 | 10.23 | 8.03 |
| change per clip: median · the clip that lost most | +0.01 · −0.03 | −0.01 · −0.10 | −0.01 · −0.08 |
MRX on the same clips: 10.92 · 5.17 · 5.72 dB (audiojs/mrx).
Model tree for audiojs/tiger-dnr
Base model
JusperLee/TIGER-DnR