Supertonic 3: int8 text encoder and vector estimator
This repo holds int8 builds of two of the four ONNX graphs in Supertone/supertonic-3. It exists for Ko TTS, an offline Korean text-to-speech engine for Android and de-Googled phones.
| File | Size | SHA-256 |
|---|---|---|
text_encoder.onnx |
16 MB | d32a22d345ecbc288b5fc121ad0aa27bd4708f9617b6c7372410836f6e02db71 |
vector_estimator.onnx |
66 MB | 9c6408bdf36ef1fa534a93baac7f828c0abb5b1b87b792153497603fcf0f5763 |
The duration predictor, vocoder, configs and voice styles are unchanged. Take them from the upstream repo at commit 724fb5abbf5502583fb520898d45929e62f02c0b.
How they were made
These files come from ONNX Runtime dynamic quantization (quantize.py):
- MatMul and Conv weights are quantized to uint8, per channel.
- Depthwise convolutions stay fp32, because ONNX Runtime's CPU kernel has no grouped
ConvInteger. - The build is deterministic, so running the script on the upstream commit reproduces the hashes above.
The vocoder is deliberately left fp32. Quantizing it turns the output into noise: 87% character error in a Whisper round-trip test.
Quality and speed
- Korean accuracy: Whisper-small round-trip character error was 4.6% with these files versus 5.3% for fp32, on 16 hospital and parenting phrases.
- Pixel 10a (Tensor G4), 4 threads, 6 steps:
| Model | First audio | Synthesis speed |
|---|---|---|
| int8 (these files) | 0.33 s | 6× faster than real time |
| fp32 | 0.70 s | 2.4× faster than real time |
License
This is a derivative of Supertonic 3 by Supertone Inc. and is released under the same BigScience OpenRAIL-M license, use restrictions included. You may not use it for impersonation, disinformation, harming people, or any of the other uses the license prohibits.
Model tree for robbiemed/supertonic-3-int8
Base model
Supertone/supertonic-3