Spark-TTS 0.5B β€” ONNX for phoonnx

ONNX weights for Spark-TTS 0.5B (SparkAudio), packaged for the phoonnx sparktts engine.

Spark-TTS is a decoder-only language model on a Qwen2.5-0.5B backbone. It does not predict audio. It predicts BiCodec tokens, and BiCodec turns those tokens into a waveform. BiCodec keeps two streams: 32 global tokens that carry the speaker, and one semantic stream at 50 Hz that carries what is said. The language model reads the text and the 32 global tokens, then writes the semantic stream.

Files

File Stage Input Output
model.onnx (+ model.onnx_data) Qwen2 language model, KV-cached input_ids, attention_mask, position_ids, past_key_values.* logits, present.*
bicodec_vocoder.onnx BiCodec decoder semantic_tokens [1,T], global_tokens [1,1,32] waveform [1,1,N] at 16 kHz
wav2vec2_model.onnx cloning front end wav [1,N] at 16 kHz, zero-mean unit-variance feat [1,T,1024]
bicodec_encoder_quantizer.onnx cloning: semantic stream feat [1,T,1024] semantic_tokens [1,T]
speaker_encoder_tokenizer.onnx cloning: speaker stream spec [1,513,T] magnitude STFT global_tokens [1,1,32]
tokenizer.json the model's own subword BPE
voices/*.json preset speakers: 32 global tokens each

The short-time Fourier transform in front of the speaker encoder is not in the graphs. ONNX has no complex dtype, so neither torch exporter can lower torch.stft. The mel filterbank projection is inside speaker_encoder_tokenizer.onnx; a caller only has to supply the magnitude spectrogram of torchaudio.transforms.Spectrogram(n_fft=1024, win_length=640, hop_length=320, power=1).

Provenance

  • model.onnx is mirrored from Fhrozen/Spark-TTS-0.5B-ONNX, verified against the SparkAudio torch weights (below).
  • the four BiCodec / wav2vec2 graphs were exported from SparkAudio/Spark-TTS-0.5B with opset 17.
  • the preset voices were minted with Spark-TTS controllable-TTS mode (gender, moderate pitch, moderate speed) and then frozen, so each voice is stable across calls.

Verified against the torch model

Stage Metric Result
Language model, prefill max absolute logit difference 1.3e-3
Language model, 8 decode steps max absolute logit difference 6.6e-4
Language model, 8 decode steps greedy token agreement 8/8
Speaker tokenizer token mismatches 0/32
Encoder + quantizer token mismatches 0/149
BiCodec vocoder max absolute waveform difference 1.2e-6
Magnitude STFT (NumPy vs torchaudio) max absolute difference 4.8e-7
wav2vec2 features (real clip) max relative difference 1.8e-5
wav2vec2 β†’ quantizer, end to end semantic token mismatches 0/199

The quantized language models published upstream (model_q4, model_q4f16, model_int8) are not mirrored here. They diverge from the torch model by tens of logits and disagree with greedy decoding, so they are not safe defaults.

License and attribution

Spark-TTS 0.5B is released by SparkAudio under CC-BY-NC-SA-4.0; the upstream code is Apache-2.0. This repository redistributes converted weights under the same terms and adds no license of its own. Cite the original work:

Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens β€” SparkAudio. https://github.com/SparkAudio/Spark-TTS

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for OpenVoiceOS/phoonnx-spark-tts

Quantized
(5)
this model

Collection including OpenVoiceOS/phoonnx-spark-tts