FlowSVE-MS: checkpoint, data manifests and benchmark audio

Anonymous release accompanying a submission under double-blind review. The code that loads the checkpoint and renders the data is in the anonymous code repository linked from the paper.

File Content
flowsve_ms_ep46.ckpt FlowSVE-MS, final checkpoint (epoch 46, 528,750 updates), EMA weights, 65.7M parameters, PyTorch Lightning format reduced to weights and hyperparameters
data/aug_train_manifest.csv.gz 720,000-pair training recipe (1:1 speech:singing, 1000 h) with every degradation parameter
data/aug_valid_manifest.csv.gz 66,876-pair singing validation recipe
benchmark/benchmark_public.tgz the two Acapella conditions of the singing benchmark (MIR-1K and OpenSinger sources): degraded inputs, clean references and the outputs of every evaluated system, FLAC 44.1 kHz stereo

Source paths inside the manifests are relative to the corpus roots (GTSinger/..., Singing/<corpus>/..., Speech/<corpus>/..., interference_pool/...); see data/README.md in the code repository for the rendering procedure. No corpus audio is redistributed; the manifests and benchmark audio are released under CC BY 4.0, the checkpoint under MIT.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support