FlowSVE-MS: checkpoint, data manifests and benchmark audio
Anonymous release accompanying a submission under double-blind review. The code that loads the checkpoint and renders the data is in the anonymous code repository linked from the paper.
| File | Content |
|---|---|
flowsve_ms_ep46.ckpt |
FlowSVE-MS, final checkpoint (epoch 46, 528,750 updates), EMA weights, 65.7M parameters, PyTorch Lightning format reduced to weights and hyperparameters |
data/aug_train_manifest.csv.gz |
720,000-pair training recipe (1:1 speech:singing, 1000 h) with every degradation parameter |
data/aug_valid_manifest.csv.gz |
66,876-pair singing validation recipe |
benchmark/benchmark_public.tgz |
the two Acapella conditions of the singing benchmark (MIR-1K and OpenSinger sources): degraded inputs, clean references and the outputs of every evaluated system, FLAC 44.1 kHz stereo |
Source paths inside the manifests are relative to the corpus roots (GTSinger/...,
Singing/<corpus>/..., Speech/<corpus>/..., interference_pool/...); see data/README.md
in the code repository for the rendering procedure. No corpus audio is redistributed;
the manifests and benchmark audio are released under CC BY 4.0, the checkpoint under MIT.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support