dubbing-demucs
The demucs checkpoint the Blaze dubbing pipeline separates speech with โ 04573f0d, the
vocals specialist from facebookresearch/demucs
(MIT). Mirrored here so a Docker build pulls a pinned file instead of whatever torch.hub
serves that day.
DEMUCS_MODEL=04573f0d in the service selects it. It is roughly 4x faster than
htdemucs_ft, and the pipeline takes the background as mix-minus-vocals rather than the sum
of the other stems, so the specialist's weak drums/bass/other cost nothing โ measured within
0.1 dB of the original background, against 8.5 dB down when summing stems.
Branches: main is what builds pin. dev is for trying a different checkpoint first.
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support