SoulX-Singer (ONNX)

An ONNX conversion of SoulX-Singer by Soul AI Lab, a zero-shot singing voice synthesis model. All weights in this repository are SoulX-Singer's; nothing here was trained or fine-tuned. The files are what the built-in vocal engine of the Trevor DAW downloads on first use; they can run anywhere ONNX Runtime does (CPU, DirectML, ...).

Files

File Contents Precision
frontend.onnx Note encoders: phoneme, note pitch, note type and f0 embeddings, four ConvNeXt-V2 blocks, expansion to mel frames (mel2note) and the flow-matching decoder's condition projection fp32
dit.onnx One estimate of the flow-matching transformer (22 Llama layers with adaptive RMSNorm on the time step): noisy mel, time and condition in, flow out fp16 weights, fp32 inputs and outputs
vocos.onnx The Vocos vocoder up to the complex spectrum (real and imaginary parts, 961 bins, hop 480 at 24 kHz) fp16 weights, fp32 inputs and outputs

The sampler (Euler steps over dit.onnx with classifier-free guidance, scale 3, rescaled by 0.75), the inverse STFT ("same" padding, Hann window, n_fft 1920), G2P and the prompt mel spectrogram are not in the graphs; Trevor runs them in Rust (crates/autodaw_vocal). There is no voice in these files: SoulX-Singer clones the voice of a short prompt recording at run time.

Changes from the original weights

  • The model is split into the three graphs above; the sampling loop and the inverse STFT stay outside.
  • dit.onnx and vocos.onnx were exported from the half-precision model, so their weights are SoulX-Singer's rounded to fp16. SoulX's explicit float32 steps (RMSNorm variance, softmax) stay float32.
  • The sinusoidal time-step embedding is cast to fp16 in the graph (the original relies on CUDA autocast for that), and the half-precision graphs take and give float32 tensors.

From the same noise, the fp32 graphs match the PyTorch model (60 dB SNR); the fp16 graphs differ from the fp32 model about as much as SoulX-Singer's own fp16 GPU inference does (log-mel difference under 0.2 dB).

License and attribution

  • SoulX-Singer code and weights: Apache License 2.0 (LICENSE), Copyright Soul AI Lab.
  • The flow-matching transformer derives from Amphion (MIT License).
  • The vocoder architecture is Vocos (MIT License).

The conversion was made for Trevor; see NOTICE for what was changed.

Usage disclaimer (from SoulX-Singer)

SoulX-Singer is intended for academic research, educational purposes and legitimate applications such as personalized vocal synthesis and assistive technologies.

  • Respect intellectual property, privacy and personal consent when generating singing content.
  • Do not use the model to impersonate individuals without authorization or to create deceptive audio.
  • The developers assume no liability for any misuse of this model.

Citation

@misc{soulxsinger,
      title={SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis},
      author={Jiale Qian and Hao Meng and Tian Zheng and Pengcheng Zhu and Haopeng Lin and Yuhang Dai and Hanke Xie and Wenxiao Cao and Ruixuan Shang and Jun Wu and Hongmei Liu and Hanlin Wen and Jian Zhao and Zhonglin Jiang and Yong Chen and Shunshun Yin and Ming Tao and Jianguo Wei and Lei Xie and Xinsheng Wang},
      year={2026},
      eprint={2602.07803},
      archivePrefix={arXiv},
      primaryClass={eess.AS},
      url={https://arxiv.org/abs/2602.07803},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pacoloco71/SoulX-Singer-ONNX

Quantized
(1)
this model

Paper for pacoloco71/SoulX-Singer-ONNX