Paradee-8M v1.0
Paradee is a small English text-to-speech model. It has 8.07M parameters and is distilled from
Kokoro-82M, and it speaks one voice, Kokoro's af_heart.
- Small. The int8 model is one 9 MB ONNX file, against 325 MB for Kokoro.
- Fast. It runs about 18x faster than real time on one CPU thread, with no GPU.
- Close to its teacher. It scores 4.41 on UTMOS (Kokoro: 4.52) and has the same word error rate (5.7%).
Paper: Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model
Code, training and the Python package: github.com/sahilmahendrakar/paradee
Samples
Five held-out sentences, read by Paradee and by its teacher. The sentences are in
samples/sentences.txt.
| # | Paradee | Kokoro-82M (teacher) |
|---|---|---|
| 1 | listen | listen |
| 2 | listen | listen |
| 3 | listen | listen |
| 4 | listen | listen |
| 5 | listen | listen |
Usage
pip install git+https://github.com/sahilmahendrakar/paradee
python -m paradee "Paradee is a small voice that runs anywhere." -o hello.wav
from paradee import Paradee, SAMPLE_RATE
import soundfile as sf
tts = Paradee()
sf.write("hello.wav", tts("Paradee is a small voice that runs anywhere."), SAMPLE_RATE)
Files
| File | What it is |
|---|---|
onnx/paradee_int8.onnx |
The whole model in one graph, weights in int8 (9.0 MB). Use this one. |
onnx/paradee.onnx |
The same graph in fp32 (37 MB). It sounds the same. |
config.json |
Kokoro's phoneme vocabulary and the sample rate |
tokenizer.json |
The same vocabulary in the format transformers.js and kokoro-js load |
pytorch/text_side.pt, pytorch/decoder.pt |
PyTorch weights for the two halves, for the training code on GitHub |
samples/ |
The audio above |
The ONNX graph goes from phoneme ids to audio. It includes the phase-locking filter described below.
- Inputs:
input_ids, int64[1, T], phoneme ids fromconfig.jsonwith a0pad token at each end, at most 512 in total. The second input isspeed, float32[1], where 1.0 is normal speed. - Output:
waveform, float32[1, samples]at 24 kHz.
Phonemes must be written the way misaki, Kokoro's own
grapheme-to-phoneme library, writes them, because that is all Paradee saw in training. In the
browser, kokoro-js phonemizes with eSpeak NG instead. There,
web/misaki.js converts
eSpeak's spelling to misaki's. Without it, Whisper mishears about 31% of words, and with it about 2%.
Results
All numbers are on 200 held-out sentences. Speed is on one CPU thread of an Apple M4 Pro.
| Model (voice) | Params | File size | Speed | UTMOS | WER |
|---|---|---|---|---|---|
| Kokoro-82M, teacher (af_heart) | 81.8M | 325 MB | 7.6x | 4.52 | 5.7% |
| Paradee (af_heart) | 8.07M | 8.45 MB | 25.0x | 4.41 | 5.7% |
| Kokoro-7M-Distill (af_msa) | 7.48M | 30.1 MB | 35.5x | 4.18 | 7.4% |
| Piper, en_US-lessac-medium | 15.7M | 63.2 MB | 15.4x | 4.36 | 8.8% |
| KittenTTS nano 0.8 (Bella) | 14.0M | 56.8 MB | 10.7x | 4.01 | 5.5% |
UTMOS is a neural network trained on human ratings. It predicts how natural a clip sounds, on a scale from 1 to 5. WER is the share of words that Whisper (base) transcribes wrongly.
The table is from the paper and measures PyTorch. The released paradee_int8.onnx scores UTMOS
4.41 on the same sentences and runs about 18x faster than real time in onnxruntime on one thread.
Hugging Face Open TTS Leaderboard
As of October 9, 2026, on Hugging Face's Open TTS Leaderboard β which scores open-source TTS models on Seed TTS Eval and CV3 Eval β Paradee-8M-v1.0 ranks:
- π₯ 2nd on English WER (1.47%), behind only its own teacher, Kokoro-82M (1.46%).
- On the Pareto front for model size vs. WER β no model of comparable size scores a lower WER.
- In the π£οΈ CPU streaming benchmark (batch size 1, single CPU thread): 1st on RTFx (11.48, i.e. ~11.5x faster than real-time) and 4th on time-to-first-audio (390 ms median).
How it was made
Paradee is Kokoro's own code at smaller widths: a 4.23M text side and a 3.85M decoder. The two halves were trained separately against the frozen teacher, on 12,000 WikiText-103 sentences (23.9 hours) read by Kokoro.
- The text side learns to predict the teacher's phoneme durations, pitch, loudness and phoneme features.
- The decoder learns to turn the teacher's saved values into the teacher's audio, first with spectrogram losses and then with adversarial training.
- The halves are joined with no further training.
- A filter with no parameters corrects the phase of voiced sound between 2 and 8 kHz. Phase is the timing of each frequency's wave. This removes a slight buzz that the small decoder otherwise leaves.
All training ran on one MacBook Pro.
Limitations
- English only, with American pronunciation, and one voice.
- Numbers, abbreviations and unusual words are pronounced only as well as misaki handles them.
- A faint buzz can still be heard on some voiced sounds, though much less than without the filter.
License
Apache 2.0, the same as Kokoro-82M.
Citation
@misc{mahendrakar2026paradee,
title = {Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model},
author = {Mahendrakar, Sahil},
year = {2026},
url = {https://huggingface.co/sahilmahendrakar/Paradee-8M-v1.0}
}
- Downloads last month
- 1,307