phoonnx β Indic Parler-TTS (ONNX)
ONNX export of ai4bharat/indic-parler-tts
for the phoonnx indic_parler engine.
Indic Parler-TTS speaks 20 Indic languages and English. You select the voice with a natural-language description ("Rohit's voice is clear and expressive..."), not with a speaker id or a reference clip.
Graphs
| File | Size | Purpose |
|---|---|---|
text_encoder.onnx |
1.4 GB | Flan-T5 encoder over the voice description |
decoder_prefill.onnx |
2.1 GB | First AR step; emits self-attention and cross-attention KV |
decoder_decode.onnx |
1.3 GB | Later AR steps; consumes both caches, emits self-attention KV only |
dac_decoder.onnx |
217 MB | DAC 44.1 kHz codec decoder (9 codebooks) |
tokenizer.json |
prompt tokenizer (the text you want spoken) | |
description_tokenizer.json |
description tokenizer (Flan-T5 vocabulary) | |
config.json |
self-describing phoonnx config (engine: indic_parler) |
The two tokenizers are different vocabularies. Upstream is explicit about this: one tokenizer for the prompt, one for the description.
Architecture
description --> Flan-T5 encoder --> encoder states --> cross-attention (all 24 layers)
prompt --> embed_prompts --> prepended to the decoder input embeddings
decoder --> 9 delayed DAC codebooks --> DAC decoder --> 44.1 kHz mono
Cross-attention keys and values do not change while decoding, so decoder_prefill
computes them once and decoder_decode reads them back unchanged.
Parity
Every graph is float32 and was checked against the PyTorch model
(parler_tts.ParlerTTSForConditionalGeneration, eager attention) on ser9 CPU:
| Check | Result |
|---|---|
| Text encoder, max abs diff (hi / en / ta) | 3.9e-07 / 6.5e-06 / 3.2e-07 |
| Prefill KV, max abs diff (self K/V, cross K/V) | 8.1e-06 / 4.5e-06 / 2.4e-06 / 1.2e-06 |
| Greedy codes, 60 steps | identical, all 3 languages |
| Waveform correlation vs torch | 0.9999999917 / 1.0000000000 / 1.0000000000 |
| Waveform RMSE vs torch | 1.3e-08 / 5.0e-08 / 1.5e-08 |
No quantised variants are published: the quantised graphs were not verified, and phoonnx does not ship unverified exports.
Notice
Upstream model: ai4bharat/indic-parler-tts by AI4Bharat,
built on huggingface/parler-tts by Yoach
Lacombe, Vaibhav Srivastav and Sanchit Gandhi. Apache-2.0, preserved from upstream.
The DAC codec is Descript's, as vendored by upstream.
The weights are unmodified: the export wraps the PyTorch modules and traces them. No surgery, no quantisation.
- Downloads last month
- -
Model tree for OpenVoiceOS/phoonnx-indic-parler
Base model
ai4bharat/indic-parler-tts