MDX23C InstVoc HQ β Core AI (float16)
Vocal / instrumental separation for Apple silicon: MDX23C-8KFFT-InstVoc_HQ
converted to a Core AI asset (stems-mdx23c-instvoc-float16.aimodel, 214 MB) for
on-device use from Swift.
Credit
- Architecture: TFC-TDF-UNet v3 (MDX23C) β Kim & Lee.
- Implementation: ZFTurbo/Music-Source-Separation-Training (MIT).
- Weights:
MDX23C-8KFFT-InstVoc_HQfrom Ultimate Vocal Remover's public model repository (MIT).
This repository only converts those weights; all credit for the model belongs to them.
What the asset computes
The network between the spectrogram and the stem spectrograms β the STFT, chunking and inverse STFT run in the host:
- input
spec[1, 4, 4096, 256]float32 β left re, left im, right re, right im; bins 0β4095 of an STFT with n_fft 8192, hop 1024, periodic Hann, centred; one chunk is 261,120 samples at 44.1 kHz. - output
stems[1, 8, 4096, 256]float32 β vocals (4 channels), then instrumental.
Weights and arithmetic are float16 behind float32 input and output.
Measured (macOS 27, Apple silicon)
| this asset | float32 export | 8-bit Core ML package | |
|---|---|---|---|
| one chunk vs PyTorch float32 | 62.6 dB | 123.9 dB | 35.3 dB |
| time per chunk (GPU) | 0.27 s | 0.31 s | 0.74 s |
| a 30 s mix, end to end | 3.0 s | 3.5 s | 6.9 s |
On three 30-second mixes with a known vocal, all three scored the same vocal and
instrumental SDR to within 0.01 dB. Repeated runs are bit-for-bit identical. The export
asserts the converted network equals upstream's model(chunk) exactly before saving.
Use
With stems, which fetches this on first use:
stems song.mp3 # song-vocal.wav + song-instrumental.wav
From Swift, with swift-vocal-isolation:
let separator = try await MDXSeparator(contentsOf: assetURL)
let stems = try await separator.separate(contentsOf: songURL)
Converted with Tools/export_mdx.py --precision float16 in swift-vocal-isolation.