Supertonic 3, repacked into one file
Supertonic 3 is Supertone's on-device text to speech model: 31 languages, about 99M parameters, 44.1 kHz. This repository holds the same model repacked into a single safetensors file, at full precision, at 8 bits and at 4 bits, for an engine that is one file of C or one file of Swift: leok7v/supertonic.tts.
THESE FILES ARE MODIFIED COPIES of the upstream release. The model is Supertone's. This repack is not affiliated with Supertone and not endorsed by them. What was changed is listed below, and the license is theirs, unchanged.
Files
| file | size | what |
|---|---|---|
supertonic-fp32.safetensors |
399 MB | every weight as released, float32 |
supertonic-q8.safetensors |
111 MB | the matrices as int8, see below |
supertonic-q4.safetensors |
79 MB | most matrices at 4 bits, see below |
LICENSE |
BigScience Open RAIL-M, the upstream file |
Each file is the whole model: the four modules (duration predictor, text encoder, vector field, vocoder), the ten voices, the character table and the Unicode tables of the text frontend. Nothing else is needed to speak.
What was changed
The source is Supertone/supertonic-3 at revision
724fb5abbf5502583fb520898d45929e62f02c0b, taken through
mlx-community/supertonic-3-mlx at
d1fb07e6ecd59ab02dcc60bc202ff13c014ef292, whose tensors are byte for
byte the initializers of the upstream ONNX files.
- The 698 float32 tensors of the four modules are stored under their own names. In the fp32 file they are unchanged, byte for byte.
- Added from the same release: the ten voices as
voice.F1.ttl,voice.F1.dpand so on toM5, the character table asindexer, andtime.freqs, a constant of the vector field graph. - Added and not Supertone's:
nfkd.offsets,nfkd.data,nfkd.classandword.edges, tables of Unicode 15.0 the text frontend needs. - In the q8 file every
.weightof 4096 elements or more is rounded: int8 codes with one float32 scale per slice of the first axis. The duration predictor and the vocoder's first convolution are float16 instead, and the character embeddings stay float32. The weights move by 0.8% to 1.3% of their norm (0.02% for float16). - In the q4 file those matrices are 4 bit codes in blocks of 32
weights, each block with its own scale and floor, 4.625 bits per
weight. Three matrices of the vocoder turn that rounding into audible
noise and are kept wider:
pwconv2andhead.layer1as in the q8 file,head.layer2as float16, with the duration predictor and the vocoder's first convolution. The 4 bit weights move by 5% to 8% of their norm. config/tts.jsonis kept verbatim in the file's metadata. Nothing reads it.
No weight was retrained. The q8 file is a different numerical model: it does not reproduce the fp32 output sample for sample. In listening tests over ten voices, in English and Russian, 8 bits everywhere could not be told from fp32, and this file is closer to fp32 than that one was. The q4 file is a larger step and the furthest of the three from fp32. Its mix was chosen by listening for what the rounding adds, a hiss above 12 kHz and noise in the pauses, and keeping wide the matrices that cause it. Take q8 unless the 32 MB matter.
The format
Both files are valid safetensors: the safetensors library, numpy and
the viewer on this page read them. They are also laid out so that an
engine can use them with one mmap and no JSON parser.
- The JSON header is padded with spaces so the data starts on a 16 KB page.
- The first tensor is
tts.index: the bytesSUPERTON, an int64 count, then 160 byte records sorted by name. A record isname[120], int32bits, an int32 spare, int32shape[4], an int64 element count and an int64 offset from the start of the file. The engine reads the header length, skips the JSON and searches this table. bits0 is the dtype as the header states it, 16 is float16, and 8 is int8 codes whose scales are the tensorNAME.scale.bits4 is a uint8 tensor[slices][bytes], and its record keeps the shape of the matrix. A slice of the first axis is blocks of 32 weights. Every eight blocks share a 20 byte head: a float16 unit of scale, a float16 unit of floor, eight uint8 steps of scale and eight uint8 steps of floor. The heads of a slice come first, then 16 bytes of codes per block, low nibble first. A weight isunit * step * code - unit * stepof the floor.- A tensor of 4096 elements or more starts on a 16 KB page, a smaller
one on a 128 byte boundary. safetensors allows no byte outside a
tensor, so the padding is itself tensors of zeros named
hole.0000and up: 24 of them in the fp32 file, 37 in the q8 file. Ignore them. - The metadata holds
layout,weights(the precision spec),source,changes,license,unicodeandconfig/tts.json.
Reading a matrix back in Python:
from safetensors.numpy import load_file
tensors = load_file("supertonic-q8.safetensors")
name = "vocoder.tts.ae.decoder.convnext.0.pwconv1.weight"
weight = tensors[name]
if weight.dtype == "int8":
scale = tensors[name + ".scale"]
weight = scale.reshape(-1, *[1] * (weight.ndim - 1)) * weight.astype("float32")
weight = weight.astype("float32")
A file loaded and saved again by a library keeps every tensor and loses the order and the alignment. The engine refuses such a file; any other reader does not care.
The engine at supertonic.tts reads all three files.
Speaking with it
git clone https://github.com/leok7v/supertonic.tts
cd supertonic.tts
HF=https://huggingface.co/leok7v/supertonic/resolve/main
curl -L -O $HF/supertonic-q8.safetensors
clang -std=c2x -O3 -ffp-contract=off -DACCELERATE -o tts tts.c \
-framework Accelerate
./tts speak --pack supertonic-q8.safetensors --voice M1 \
--lang en --text "Hello world." --out hello.wav
Without -DACCELERATE and the framework the same file builds with libc
alone, about ten times slower. tts.swift is the same engine, function
for function, and writes the same WAV byte for byte.
Measured on an Apple M3, one thread, 8 flow steps, the q8 file, the Accelerate build:
| "Hello world." | 284 characters | |
|---|---|---|
| audio | 1.35 s | 16.8 s |
| wall clock | 0.21 s | 1.50 s |
| peak private memory | 9 MB | 28 MB |
| resident, the file included | 112 MB | 131 MB |
With the q4 file the same two texts take 0.28 s and 1.56 s, and 80 MB and 99 MB resident: 70 ms more per chunk of text, 32 MB less.
The engine is checked stage by stage against the upstream ONNX SDK: with the fp32 file, 497 named taps per text agree to float32 noise, for six texts in five languages.
The model
From the upstream card. Supertonic is a lightweight text to speech system for local inference, with no cloud call required for synthesis. Supertonic 3 expands the open weight release from 5 to 31 languages, improves reading stability and reduces repeat and skip failures. The upstream card has the accuracy, speed and size measurements, a demo, and the Python SDK, which reads the ONNX files of the upstream repository, not the files here.
Voices: F1 to F5 and M1 to M5.
| Code | Language | Code | Language | Code | Language | Code | Language |
|---|---|---|---|---|---|---|---|
en |
English | ko |
Korean | ja |
Japanese | ar |
Arabic |
bg |
Bulgarian | cs |
Czech | da |
Danish | de |
German |
el |
Greek | es |
Spanish | et |
Estonian | fi |
Finnish |
fr |
French | hi |
Hindi | hr |
Croatian | hu |
Hungarian |
id |
Indonesian | it |
Italian | lt |
Lithuanian | lv |
Latvian |
nl |
Dutch | pl |
Polish | pt |
Portuguese | ro |
Romanian |
ru |
Russian | sk |
Slovak | sl |
Slovenian | sv |
Swedish |
tr |
Turkish | uk |
Ukrainian | vi |
Vietnamese |
License
The model is released by Supertone under the BigScience Open RAIL-M License, and so are these files, which are derivatives of it. The LICENSE file in this repository is the upstream one, unchanged, and it is the text that binds; what follows is a summary.
- You may use, copy, modify and redistribute the files, also commercially.
- Anyone you pass them to must get a copy of the license, and the use restrictions of its Attachment A must stay an enforceable condition on them and on everything derived from them.
- Those restrictions forbid, among other things, impersonating a person without their consent, defaming or harassing others, and publishing generated content without disclosing that a machine made it. Read Attachment A before you ship a voice.
- Modified files must say that they were modified. These were, as described above.
Copyright (c) 2026 Supertone Inc. The Unicode tables are derived from the Unicode Character Database. Supertone's sample code is MIT licensed and is not part of this repository.
Model tree for leok7v/supertonic
Base model
Supertone/supertonic-3