Supertonic 3 β mobile-optimized for sherpa-onnx
Supertonic 3 (Supertone, 2026-04, 31 languages)
repacked for fast on-device CPU inference with
sherpa-onnx (OfflineTtsSupertonicModelConfig).
Used by the Translaty app for fully offline speech.
What changed
Only vector_estimator.int8.onnx, the flow-matching network that runs once per
denoising step and dominates synthesis time. In the upstream graph ~85 % of that
time is spent in point-wise Conv1d (kernel 1) layers that stay in fp32 even in
the int8 export. conv2mm.py rewrites each of them as Transpose β MatMul β Transpose
and applies per-channel dynamic int8 quantization to the MatMuls, so ONNX Runtime
uses its integer GEMM kernels. The output matches the fp32 model more closely than
the upstream int8 export (relative error 4.2 % vs 4.5 % on one step).
Every other file is unchanged: duration_predictor.onnx and text_encoder.onnx
are the upstream fp32 files, vocoder.int8.onnx, tts.json, unicode_indexer.bin
and voice.bin come from sherpa-onnx's sherpa-onnx-supertonic-3-tts-int8-2026-05-11.
Quantizing the vocoder the same way was tried and rejected: it cost ~0.6 UTMOS.
voice.bin holds 10 voices: sid 0β4 = F1βF5, sid 5β9 = M1βM5.
Pass the language as extra: {"lang": "<code>"}.
Measured (Galaxy Note 8, Exynos 8895, 2017, 4 threads)
| Model | Steps | RTF | UTMOS |
|---|---|---|---|
| upstream sherpa-onnx int8 | 8 | 1.65 | 4.52 |
| upstream sherpa-onnx int8 | 4 | 0.87 | 3.80 |
| this repo | 4 | 0.39 | 4.42 |
| this repo | 8 | 0.67 | 4.51 |
RTF = synthesis time / audio duration (lower is faster). UTMOS22 on an English sentence. Whisper large-v3-turbo round-trips all 14 tested languages correctly.
License
OpenRAIL-M, inherited from Supertonic 3 β see LICENSE, including its use restrictions.
Model tree for WilburDev/supertonic-3-mobile
Base model
Supertone/supertonic-3