SCNet-large, ONNX (int8 weights)

SCNet-large (Tong, Zhu, Chen, Kang, Jiang, Li, Wu, Meng, "SCNet: Sparse Compression Network for Music Source Separation", ICASSP 2024, arXiv:2401.13276): band-split convolutions around dual-path LSTMs on the complex spectrogram, 41.2 M parameters, trained by its author on MUSDB18-HQ. It splits a song into four stems: drums, bass, other, vocals. This is the network between its STFT and iSTFT, as @audio/neural-separate runs it (model: 'scnet-large').

File Size SHA-256
scnet-large.int8.onnx 45.1 MB (45,082,486 bytes) b2dc586a1e0e6c0afe4915e9057ea29111397b7bbf589edc3d85de91de79cd72

The float32 export it is made from is 169.2 MB; float16 weights would be 86.5 MB.

Source

  • Model and code: starrytong/SCNet (MIT), at 5d95bf96b19c3eede63248d171efeca8e3abb948.
  • Checkpoint: SCNet-large_starrytong_fixed.ckpt (SHA-256 65900dfa07d6b6e5d784c0f143920200a4bd281d6e78a806c549d0b912d5885e), release v1.0.9 of ZFTurbo/Music-Source-Separation-Training (MIT), with its config_musdb18_scnet_large_starrytong.yaml; the model code that repository's models/scnet at 84b1eac0887756b4f1a9d7a1ff49105939749ed2.

Licence and attribution

MIT, Copyright (c) 2024 starrytong (LICENSE). The weights' author, in starrytong/SCNet#35 (2026-07-17):

I confirm that the released SCNet and SCNet-large pretrained weights are distributed under the MIT License, consistent with the source code. You are welcome to redistribute the original checkpoints and format-converted versions, including ONNX exports, as part of your MIT-licensed tool, with appropriate attribution.

SCNet-large by its authors (starrytong/SCNet); the checkpoint as Music-Source-Separation-Training (Roman Solovyev) distributes it; ONNX export and compaction by audiojs. Trained on MUSDB18-HQ (Rafii et al., 2019), licensed for educational use; whether that reaches the weights no project has settled.

@inproceedings{tong2024scnet,
  title     = {SCNet: Sparse Compression Network for Music Source Separation},
  author    = {Tong, Weinan and Zhu, Jiaxu and Chen, Jun and Kang, Shiyin and Jiang, Tao and Li, Yang and Wu, Zhiyong and Meng, Helen},
  booktitle = {IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
  year      = {2024},
  eprint    = {2401.13276},
  archivePrefix = {arXiv}
}

Graph

One 11 s segment (485,100 samples at 44.1 kHz, padded to 476 frames) per run:

name shape
input mix_spec [1, 4, 2049, 476] STFT: n 4096, hop 1024, no window, scaled by 1/√4096, centered with reflect padding; L re, L im, R re, R im
output stems_spec [1, 16, 2049, 476] each source (drums, bass, other, vocals), channel, re and im

The segments (every 2.75 s), their fades and the input's normalization follow Music-Source-Separation-Training's demix(); @audio/neural-separate's README, Algorithm, has them.

import separate from '@audio/neural-separate'
let { stems } = await separate([left, right], { sampleRate: 44100, model: 'scnet-large' })

Export, compaction, verification

scripts/export-scnet.py --model scnet-large --verify exports the network (its rFFT over time as cosine and sine products, its GroupNorm statistics reduced axis by axis) and compares the graph with SCNet.forward on noise and tones: max |diff| ≀ 2.8e-6 of max |y|; the package's pipeline matches SCNet.forward on its segments to 114–134 dB SNR per stem. scripts/compact.py --model scnet-large --calibrate <two MUSDB18 training previews> makes this file from it:

  • Weights: 83 of 88 stored in int8 (99.2 % of the values; symmetric, a scale per output channel, an LSTM's per gate row and direction), the five layers ending the decoder in float16 (decoder.2.0's convolution, decoder.1.1's three transposed convolutions, decoder.2.1's first): rounded alone to int8, each moves the output 27 to 38 dB under its power; all 88 together, 24.0 dB.
  • Computed in float32 on every backend: each weight is Cast and multiplied by its scale in the graph, which onnxruntime folds at load (the session holds float32 weights).
  • Folded and named short, changing no value: with the input's shape fixed, every value computable from the weights and the shapes alone is stored as the graph computes it (its DFT matrices, made in float64 from a Range, which onnxruntime-web's WebGPU session cannot place); node and value names are base-36 counters.
  • Against the export, on the calibration previews: max |diff| 7.1e-3 of max |y|, SNR 43.7 dB.

Quality

The 50 MUSDB18 test previews, BSSEval v4 SDR (museval), the median over songs, dB:

vocals drums bass other
export (float32) 11.00 10.27 8.21 6.87
this file 10.96 10.26 8.20 6.92
change per song: median Β· the song that lost most βˆ’0.00 Β· βˆ’0.43 βˆ’0.00 Β· βˆ’0.03 βˆ’0.00 Β· βˆ’0.04 +0.00 Β· βˆ’0.07

The βˆ’0.43 dB is a song whose vocal stem is near silence (PR - Happy Daze, βˆ’1.9 dB SDR as exported). Remixes (the input plus (g βˆ’ 1) times a stem, against the true remix): vocals +6 dB 20.21 β†’ 20.22, vocals βˆ’6 dB 22.11 β†’ 22.09, drums βˆ’6 dB 21.88 β†’ 21.89. Float16 weights (86.5 MB) change no median by more than 0.002 dB.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for audiojs/scnet-large