MRX, ONNX (int8 weights)
MRX (Petermann, Wichern, Wang, Le Roux, "The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World
Soundtracks", ICASSP 2022, arXiv:2110.09958): magnitudes at three STFT resolutions
into one hidden layer, a BLSTM per source, a real mask per source and resolution; 30.5 M parameters, MERL's own
checkpoint trained on Divide and Remaster with the SNR loss. It splits a soundtrack into three stems: dialogue, music,
effects. This is the network between its STFTs and iSTFTs, as
@audio/neural-separate runs it (model: 'mrx').
| File | Size | SHA-256 |
|---|---|---|
mrx.int8.onnx |
31.3 MB (31,287,081 bytes) | d876c92d224f5d8cd2207d528b23369da895a6de71fb24132416fcb16f9f8424 |
The float32 export it is made from is 122.3 MB; float16 weights would be 61.5 MB.
Source
- Model, code and checkpoint: merlresearch/cocktail-fork-separation
at
19b3de827ebc4bfb014570cf92dd32b4ee3b6921:mrx.py,checkpoints/default_mrx_pre_trained_weights.pth.
Licence and attribution
MIT, Copyright (c) 2023 Mitsubishi Electric Research Laboratories (MERL) (LICENSE). The repository's
.reuse/dep5
names the checkpoints:
Files: checkpoints/default_mrx_pre_trained_weights.pth [and its three other checkpoints] Copyright: 2023 Mitsubishi Electric Research Laboratories (MERL) License: MIT
MRX by Darius Petermann, Gordon Wichern, Zhong-Qiu Wang and Jonathan Le Roux (MERL); ONNX export and compaction by audiojs. Trained on Divide and Remaster v2, whose music and effects come from FMA and FSD50K clips, each under its own licence, some non-commercial.
@inproceedings{petermann2022cocktail,
title = {The Cocktail Fork Problem: Three-Stem Audio Separation for Real-World Soundtracks},
author = {Petermann, Darius and Wichern, Gordon and Wang, Zhong-Qiu and Le Roux, Jonathan},
booktitle = {IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
year = {2022},
eprint = {2110.09958},
archivePrefix = {arXiv}
}
Graph
Each channel apart (B), any length (T frames):
| name | shape | ||
|---|---|---|---|
| input | mag_1024, mag_2048, mag_8192 |
[B, 513, T], [B, 1025, T], [B, 4097, T] | magnitude STFTs: Hann windows of 1024, 2048, 8192, hop 256, scaled by 1/√n, centered with reflect padding, at 44.1 kHz |
| output | mask_1024, mask_2048, mask_8192 |
[B, 3, F, T] | a real mask per source (music, speech, sfx) and resolution |
A source is the sum over resolutions of the iSTFT of its mask times that resolution's spectrogram; the input is set to
-27 LUFS first and the stems scaled back (upstream's separate.py); @audio/neural-separate runs it in 20 s chunks.
import separate from '@audio/neural-separate'
let { stems } = await separate([mono], { sampleRate: 48000, model: 'mrx' }) // stems.dialogue, .music, .effects
Export, compaction, verification
scripts/export-mrx.py --verify exports the network and compares the sources rebuilt from the graph with
MRX.forward on noise and tones: max |diff| ≤ 1.2e-6 of max |y|; the package's pipeline matches upstream's
separate_soundtrack to 101–134 dB SNR per stem. scripts/compact.py --model mrx --calibrate <a Divide and Remaster v3 tuning clip> makes this file from it:
- Weights: all 39 stored in int8 (symmetric, a scale per output channel, an LSTM's per gate row and direction): rounded together they move the masked spectra 40.4 dB under their power on the calibration clip, within the 40 dB the script allows before keeping any in float16.
- Computed in float32 on every backend: each weight is Cast and multiplied by its scale in the graph, which onnxruntime folds at load (the session holds float32 weights).
- Node and value names shortened (its input length is free, so nothing is folded).
- Against the export on the calibration clip: max |diff| 9.9e-3 of max |y|, SNR 40.4 dB (the worst resolution).
Quality
30 clips of Divide and Remaster v3's English test set (every 40th of its 1200 clips of 60 s; Watcharasupat, Wu, Orife, 2024, CC BY-SA 4.0): each stem against its reference over the whole clip, SNR, the median over clips, dB:
| dialogue | music | effects | |
|---|---|---|---|
| export (float32) | 10.92 | 5.17 | 5.72 |
| this file | 10.92 | 5.17 | 5.70 |
| change per clip: median · the clip that lost most | +0.00 · −0.01 | +0.00 · −0.04 | +0.00 · −0.02 |
SI-SDR (mean over clips) 9.819 · 2.804 · 3.275 against the export's 9.815 · 2.805 · 3.273. Remixes (a stem 6 dB up or down, against the true remix): dialogue +6 dB, music −6 dB, effects −6 dB each within 0.03 dB of the export's median, no clip lower by more than 0.04 dB. Float16 weights (61.5 MB) change no median by more than 0.001 dB.