Supertonic 3, repacked into one file

Supertonic 3 is Supertone's on-device text to speech model: 31 languages, about 99M parameters, 44.1 kHz. This repository holds the same model repacked into a single safetensors file, at full precision, at 8 bits and at 4 bits, for an engine that is one file of C or one file of Swift: leok7v/supertonic.tts.

THESE FILES ARE MODIFIED COPIES of the upstream release. The model is Supertone's. This repack is not affiliated with Supertone and not endorsed by them. What was changed is listed below, and the license is theirs, unchanged.

Files

file size what
supertonic-fp32.safetensors 399 MB every weight as released, float32
supertonic-q8.safetensors 111 MB the matrices as int8, see below
supertonic-q4.safetensors 79 MB most matrices at 4 bits, see below
LICENSE BigScience Open RAIL-M, the upstream file

Each file is the whole model: the four modules (duration predictor, text encoder, vector field, vocoder), the ten voices, the character table and the Unicode tables of the text frontend. Nothing else is needed to speak.

What was changed

The source is Supertone/supertonic-3 at revision 724fb5abbf5502583fb520898d45929e62f02c0b, taken through mlx-community/supertonic-3-mlx at d1fb07e6ecd59ab02dcc60bc202ff13c014ef292, whose tensors are byte for byte the initializers of the upstream ONNX files.

  • The 698 float32 tensors of the four modules are stored under their own names. In the fp32 file they are unchanged, byte for byte.
  • Added from the same release: the ten voices as voice.F1.ttl, voice.F1.dp and so on to M5, the character table as indexer, and time.freqs, a constant of the vector field graph.
  • Added and not Supertone's: nfkd.offsets, nfkd.data, nfkd.class and word.edges, tables of Unicode 15.0 the text frontend needs.
  • In the q8 file every .weight of 4096 elements or more is rounded: int8 codes with one float32 scale per slice of the first axis. The duration predictor and the vocoder's first convolution are float16 instead, and the character embeddings stay float32. The weights move by 0.8% to 1.3% of their norm (0.02% for float16).
  • In the q4 file those matrices are 4 bit codes in blocks of 32 weights, each block with its own scale and floor, 4.625 bits per weight. Three matrices of the vocoder turn that rounding into audible noise and are kept wider: pwconv2 and head.layer1 as in the q8 file, head.layer2 as float16, with the duration predictor and the vocoder's first convolution. The 4 bit weights move by 5% to 8% of their norm.
  • config/tts.json is kept verbatim in the file's metadata. Nothing reads it.

No weight was retrained. The q8 file is a different numerical model: it does not reproduce the fp32 output sample for sample. In listening tests over ten voices, in English and Russian, 8 bits everywhere could not be told from fp32, and this file is closer to fp32 than that one was. The q4 file is a larger step and the furthest of the three from fp32. Its mix was chosen by listening for what the rounding adds, a hiss above 12 kHz and noise in the pauses, and keeping wide the matrices that cause it. Take q8 unless the 32 MB matter.

The format

Both files are valid safetensors: the safetensors library, numpy and the viewer on this page read them. They are also laid out so that an engine can use them with one mmap and no JSON parser.

  • The JSON header is padded with spaces so the data starts on a 16 KB page.
  • The first tensor is tts.index: the bytes SUPERTON, an int64 count, then 160 byte records sorted by name. A record is name[120], int32 bits, an int32 spare, int32 shape[4], an int64 element count and an int64 offset from the start of the file. The engine reads the header length, skips the JSON and searches this table.
  • bits 0 is the dtype as the header states it, 16 is float16, and 8 is int8 codes whose scales are the tensor NAME.scale.
  • bits 4 is a uint8 tensor [slices][bytes], and its record keeps the shape of the matrix. A slice of the first axis is blocks of 32 weights. Every eight blocks share a 20 byte head: a float16 unit of scale, a float16 unit of floor, eight uint8 steps of scale and eight uint8 steps of floor. The heads of a slice come first, then 16 bytes of codes per block, low nibble first. A weight is unit * step * code - unit * step of the floor.
  • A tensor of 4096 elements or more starts on a 16 KB page, a smaller one on a 128 byte boundary. safetensors allows no byte outside a tensor, so the padding is itself tensors of zeros named hole.0000 and up: 24 of them in the fp32 file, 37 in the q8 file. Ignore them.
  • The metadata holds layout, weights (the precision spec), source, changes, license, unicode and config/tts.json.

Reading a matrix back in Python:

from safetensors.numpy import load_file

tensors = load_file("supertonic-q8.safetensors")
name = "vocoder.tts.ae.decoder.convnext.0.pwconv1.weight"
weight = tensors[name]
if weight.dtype == "int8":
    scale = tensors[name + ".scale"]
    weight = scale.reshape(-1, *[1] * (weight.ndim - 1)) * weight.astype("float32")
weight = weight.astype("float32")

A file loaded and saved again by a library keeps every tensor and loses the order and the alignment. The engine refuses such a file; any other reader does not care.

The engine at supertonic.tts reads all three files.

Speaking with it

git clone https://github.com/leok7v/supertonic.tts
cd supertonic.tts
HF=https://huggingface.co/leok7v/supertonic/resolve/main
curl -L -O $HF/supertonic-q8.safetensors
clang -std=c2x -O3 -ffp-contract=off -DACCELERATE -o tts tts.c \
      -framework Accelerate
./tts speak --pack supertonic-q8.safetensors --voice M1 \
      --lang en --text "Hello world." --out hello.wav

Without -DACCELERATE and the framework the same file builds with libc alone, about ten times slower. tts.swift is the same engine, function for function, and writes the same WAV byte for byte.

Measured on an Apple M3, one thread, 8 flow steps, the q8 file, the Accelerate build:

"Hello world." 284 characters
audio 1.35 s 16.8 s
wall clock 0.21 s 1.50 s
peak private memory 9 MB 28 MB
resident, the file included 112 MB 131 MB

With the q4 file the same two texts take 0.28 s and 1.56 s, and 80 MB and 99 MB resident: 70 ms more per chunk of text, 32 MB less.

The engine is checked stage by stage against the upstream ONNX SDK: with the fp32 file, 497 named taps per text agree to float32 noise, for six texts in five languages.

The model

From the upstream card. Supertonic is a lightweight text to speech system for local inference, with no cloud call required for synthesis. Supertonic 3 expands the open weight release from 5 to 31 languages, improves reading stability and reduces repeat and skip failures. The upstream card has the accuracy, speed and size measurements, a demo, and the Python SDK, which reads the ONNX files of the upstream repository, not the files here.

Voices: F1 to F5 and M1 to M5.

Code Language Code Language Code Language Code Language
en English ko Korean ja Japanese ar Arabic
bg Bulgarian cs Czech da Danish de German
el Greek es Spanish et Estonian fi Finnish
fr French hi Hindi hr Croatian hu Hungarian
id Indonesian it Italian lt Lithuanian lv Latvian
nl Dutch pl Polish pt Portuguese ro Romanian
ru Russian sk Slovak sl Slovenian sv Swedish
tr Turkish uk Ukrainian vi Vietnamese

License

The model is released by Supertone under the BigScience Open RAIL-M License, and so are these files, which are derivatives of it. The LICENSE file in this repository is the upstream one, unchanged, and it is the text that binds; what follows is a summary.

  • You may use, copy, modify and redistribute the files, also commercially.
  • Anyone you pass them to must get a copy of the license, and the use restrictions of its Attachment A must stay an enforceable condition on them and on everything derived from them.
  • Those restrictions forbid, among other things, impersonating a person without their consent, defaming or harassing others, and publishing generated content without disclosing that a machine made it. Read Attachment A before you ship a voice.
  • Modified files must say that they were modified. These were, as described above.

Copyright (c) 2026 Supertone Inc. The Unicode tables are derived from the Unicode Character Database. Supertone's sample code is MIT licensed and is not part of this repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for leok7v/supertonic

Quantized
(25)
this model