Supertonic 3 β€” mobile-optimized for sherpa-onnx

Supertonic 3 (Supertone, 2026-04, 31 languages) repacked for fast on-device CPU inference with sherpa-onnx (OfflineTtsSupertonicModelConfig). Used by the Translaty app for fully offline speech.

What changed

Only vector_estimator.int8.onnx, the flow-matching network that runs once per denoising step and dominates synthesis time. In the upstream graph ~85 % of that time is spent in point-wise Conv1d (kernel 1) layers that stay in fp32 even in the int8 export. conv2mm.py rewrites each of them as Transpose β†’ MatMul β†’ Transpose and applies per-channel dynamic int8 quantization to the MatMuls, so ONNX Runtime uses its integer GEMM kernels. The output matches the fp32 model more closely than the upstream int8 export (relative error 4.2 % vs 4.5 % on one step).

Every other file is unchanged: duration_predictor.onnx and text_encoder.onnx are the upstream fp32 files, vocoder.int8.onnx, tts.json, unicode_indexer.bin and voice.bin come from sherpa-onnx's sherpa-onnx-supertonic-3-tts-int8-2026-05-11. Quantizing the vocoder the same way was tried and rejected: it cost ~0.6 UTMOS.

voice.bin holds 10 voices: sid 0–4 = F1–F5, sid 5–9 = M1–M5. Pass the language as extra: {"lang": "<code>"}.

Measured (Galaxy Note 8, Exynos 8895, 2017, 4 threads)

Model Steps RTF UTMOS
upstream sherpa-onnx int8 8 1.65 4.52
upstream sherpa-onnx int8 4 0.87 3.80
this repo 4 0.39 4.42
this repo 8 0.67 4.51

RTF = synthesis time / audio duration (lower is faster). UTMOS22 on an English sentence. Whisper large-v3-turbo round-trips all 14 tested languages correctly.

License

OpenRAIL-M, inherited from Supertonic 3 β€” see LICENSE, including its use restrictions.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for WilburDev/supertonic-3-mobile

Quantized
(24)
this model