Supertonic 3: int8 text encoder and vector estimator

This repo holds int8 builds of two of the four ONNX graphs in Supertone/supertonic-3. It exists for Ko TTS, an offline Korean text-to-speech engine for Android and de-Googled phones.

File Size SHA-256
text_encoder.onnx 16 MB d32a22d345ecbc288b5fc121ad0aa27bd4708f9617b6c7372410836f6e02db71
vector_estimator.onnx 66 MB 9c6408bdf36ef1fa534a93baac7f828c0abb5b1b87b792153497603fcf0f5763

The duration predictor, vocoder, configs and voice styles are unchanged. Take them from the upstream repo at commit 724fb5abbf5502583fb520898d45929e62f02c0b.

How they were made

These files come from ONNX Runtime dynamic quantization (quantize.py):

  • MatMul and Conv weights are quantized to uint8, per channel.
  • Depthwise convolutions stay fp32, because ONNX Runtime's CPU kernel has no grouped ConvInteger.
  • The build is deterministic, so running the script on the upstream commit reproduces the hashes above.

The vocoder is deliberately left fp32. Quantizing it turns the output into noise: 87% character error in a Whisper round-trip test.

Quality and speed

  • Korean accuracy: Whisper-small round-trip character error was 4.6% with these files versus 5.3% for fp32, on 16 hospital and parenting phrases.
  • Pixel 10a (Tensor G4), 4 threads, 6 steps:
Model First audio Synthesis speed
int8 (these files) 0.33 s 6× faster than real time
fp32 0.70 s 2.4× faster than real time

License

This is a derivative of Supertonic 3 by Supertone Inc. and is released under the same BigScience OpenRAIL-M license, use restrictions included. You may not use it for impersonation, disinformation, harming people, or any of the other uses the license prohibits.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for robbiemed/supertonic-3-int8

Quantized
(24)
this model