TIGER for Divide and Remaster, ONNX (int8 weights)

TIGER (Xu, Li, Chen, Hu, "TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation", ICLR 2025, arXiv:2410.01469), its Divide and Remaster model TIGERDNR: three band-split models of 1.4 M parameters (57 bands, multi-scale convolutions, frame and frequency attention), each separating three sources and keeping one. It splits a soundtrack into three stems: dialogue, music, effects. This is the three networks between their STFT and iSTFT as one graph, as @audio/neural-separate runs it (model: 'tiger'), converted from JusperLee/TIGER-DnR.

File Size SHA-256
tiger.int8.onnx 7.3 MB (7,320,735 bytes) 2189859ac7e89cf4d9d284c6f393e737aaa2b3233ed7ee424ad622012bce7bca

The float32 export it is made from is 28.6 MB, of which 11.7 MB were node names; float16 weights would be 10.9 MB.

Source

  • Code: JusperLee/TIGER at 9f18d4a10a7137e1ce8052cfb62215179f1287b6: look2hear/models/tiger_dnr.py and the layer modules it reads.
  • Weights: JusperLee/TIGER-DnR at revision b7a59560bbca10febbcd46fb01600f868e587f57, model.safetensors (SHA-256 dd1c696e72f6adea0085ef1af640882a8260519ad666422835e387a5b4abdd2a).

Licence and attribution

Apache License 2.0 (LICENSE): the weights' model card states license: apache-2.0 (the code repository's LICENSE is MIT, its README badge Apache 2.0). Modified from the original weights: converted to ONNX with the three models in one graph; F.adaptive_avg_pool1d exported as a cumulative sum read at torch's windows, UConvBlock's global sum started at its first term and each GroupNorm's statistics reduced axis by axis (none changing the function); the weights stored in 8 bits (below).

TIGER by Mohan Xu, Kai Li, Guo Chen and Xiaolin Hu (Tsinghua University); ONNX export and compaction by audiojs. Trained on Divide and Remaster v1, sourced as v2 is (FMA and FSD50K clips, each under its own licence).

@inproceedings{xu2025tiger,
  title     = {TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation},
  author    = {Xu, Mohan and Li, Kai and Chen, Guo and Hu, Xiaolin},
  booktitle = {International Conference on Learning Representations (ICLR)},
  year      = {2025},
  eprint    = {2410.01469},
  archivePrefix = {arXiv}
}

Graph

One mono 12 s segment (529,200 samples at 44.1 kHz, 1034 frames) per run, the three models in one graph:

name shape
input mix_spec [1, 2, 1025, 1034] STFT: n 2048, hop 512, Hann, unnormalized, centered with reflect padding; re, im
output stems_spec [1, 6, 1025, 1034] dialogue, music, effects, each re and im

Segments start every 4 s, unweighted, zero-padded past the ends, their sum divided by three (TIGERDNR.wav_chunk_inference).

import separate from '@audio/neural-separate'
let { stems } = await separate([mono], { sampleRate: 48000, model: 'tiger' })   // stems.dialogue, .music, .effects

Export, compaction, verification

scripts/export-tiger.py --verify exports the graph and compares it with the three TIGER.forward on one segment of tones and noise: max |diff| ≤ 4.2e-7 of max |y|; the package's pipeline matches TIGERDNR.wav_chunk_inference to 95–140 dB SNR per stem. scripts/compact.py --model tiger --calibrate <a Divide and Remaster v3 tuning clip> makes this file from it:

  • Weights: all 444 convolution weights of over 1024 values stored in int8 (symmetric, a scale per output channel): rounded together they move the output 41.4 dB under its power on the calibration clip, within the 40 dB the script allows before keeping any in float16.
  • Computed in float32 on every backend: each weight is Cast and multiplied by its scale in the graph, which onnxruntime folds at load (the session holds float32 weights).
  • Folded and named short, changing no value: with the input's shape fixed, every value computable from the weights and the shapes alone is stored as the graph computes it (56,652 nodes to 38,817); node and value names are base-36 counters (the graph without its weights 11.7 MB to 1.9 MB).
  • Against the export on the calibration clip: max |diff| 1.7e-2 of max |y|, SNR 41.4 dB.

Quality

The 30 test clips @audio/neural-separate measures TIGER on (Divide and Remaster v3's English test set, every 40th of its 1200 clips of 60 s; Watcharasupat, Wu, Orife, 2024, CC BY-SA 4.0): each stem against its reference over the whole clip, SNR, the median over clips, dB:

dialogue music effects
export (float32) 12.68 10.24 8.06
this file 12.66 10.23 8.03
change per clip: median · the clip that lost most +0.01 · −0.03 −0.01 · −0.10 −0.01 · −0.08

MRX on the same clips: 10.92 · 5.17 · 5.72 dB (audiojs/mrx).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for audiojs/tiger-dnr

Quantized
(2)
this model

Paper for audiojs/tiger-dnr