MRX, ONNX (int8 weights)

MRX (Petermann, Wichern, Wang, Le Roux, "The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks", ICASSP 2022, arXiv:2110.09958): magnitudes at three STFT resolutions into one hidden layer, a BLSTM per source, a real mask per source and resolution; 30.5 M parameters, MERL's own checkpoint trained on Divide and Remaster with the SNR loss. It splits a soundtrack into three stems: dialogue, music, effects. This is the network between its STFTs and iSTFTs, as @audio/neural-separate runs it (model: 'mrx').

File Size SHA-256
mrx.int8.onnx 31.3 MB (31,287,081 bytes) d876c92d224f5d8cd2207d528b23369da895a6de71fb24132416fcb16f9f8424

The float32 export it is made from is 122.3 MB; float16 weights would be 61.5 MB.

Source

Licence and attribution

MIT, Copyright (c) 2023 Mitsubishi Electric Research Laboratories (MERL) (LICENSE). The repository's .reuse/dep5 names the checkpoints:

Files: checkpoints/default_mrx_pre_trained_weights.pth [and its three other checkpoints] Copyright: 2023 Mitsubishi Electric Research Laboratories (MERL) License: MIT

MRX by Darius Petermann, Gordon Wichern, Zhong-Qiu Wang and Jonathan Le Roux (MERL); ONNX export and compaction by audiojs. Trained on Divide and Remaster v2, whose music and effects come from FMA and FSD50K clips, each under its own licence, some non-commercial.

@inproceedings{petermann2022cocktail,
  title     = {The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks},
  author    = {Petermann, Darius and Wichern, Gordon and Wang, Zhong-Qiu and Le Roux, Jonathan},
  booktitle = {IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
  year      = {2022},
  eprint    = {2110.09958},
  archivePrefix = {arXiv}
}

Graph

Each channel apart (B), any length (T frames):

name shape
input mag_1024, mag_2048, mag_8192 [B, 513, T], [B, 1025, T], [B, 4097, T] magnitude STFTs: Hann windows of 1024, 2048, 8192, hop 256, scaled by 1/√n, centered with reflect padding, at 44.1 kHz
output mask_1024, mask_2048, mask_8192 [B, 3, F, T] a real mask per source (music, speech, sfx) and resolution

A source is the sum over resolutions of the iSTFT of its mask times that resolution's spectrogram; the input is set to -27 LUFS first and the stems scaled back (upstream's separate.py); @audio/neural-separate runs it in 20 s chunks.

import separate from '@audio/neural-separate'
let { stems } = await separate([mono], { sampleRate: 48000, model: 'mrx' })   // stems.dialogue, .music, .effects

Export, compaction, verification

scripts/export-mrx.py --verify exports the network and compares the sources rebuilt from the graph with MRX.forward on noise and tones: max |diff| ≤ 1.2e-6 of max |y|; the package's pipeline matches upstream's separate_soundtrack to 101–134 dB SNR per stem. scripts/compact.py --model mrx --calibrate <a Divide and Remaster v3 tuning clip> makes this file from it:

  • Weights: all 39 stored in int8 (symmetric, a scale per output channel, an LSTM's per gate row and direction): rounded together they move the masked spectra 40.4 dB under their power on the calibration clip, within the 40 dB the script allows before keeping any in float16.
  • Computed in float32 on every backend: each weight is Cast and multiplied by its scale in the graph, which onnxruntime folds at load (the session holds float32 weights).
  • Node and value names shortened (its input length is free, so nothing is folded).
  • Against the export on the calibration clip: max |diff| 9.9e-3 of max |y|, SNR 40.4 dB (the worst resolution).

Quality

30 clips of Divide and Remaster v3's English test set (every 40th of its 1200 clips of 60 s; Watcharasupat, Wu, Orife, 2024, CC BY-SA 4.0): each stem against its reference over the whole clip, SNR, the median over clips, dB:

dialogue music effects
export (float32) 10.92 5.17 5.72
this file 10.92 5.17 5.70
change per clip: median · the clip that lost most +0.00 · −0.01 +0.00 · −0.04 +0.00 · −0.02

SI-SDR (mean over clips) 9.819 · 2.804 · 3.275 against the export's 9.815 · 2.805 · 3.273. Remixes (a stem 6 dB up or down, against the true remix): dialogue +6 dB, music −6 dB, effects −6 dB each within 0.03 dB of the export's median, no clip lower by more than 0.04 dB. Float16 weights (61.5 MB) change no median by more than 0.001 dB.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for audiojs/mrx