Kokoro-7M-Distill ONNX

ONNX export of oddadmix/Kokoro-7M-Distill, an English TTS model with 7.48M parameters distilled from Kokoro-82M. It runs with onnxruntime alone, without PyTorch.

file precision size
onnx/model.onnx fp32 30.7 MB
onnx/model_fp16.onnx fp16 weights, fp32 I/O 16.1 MB
voices/af_msa.bin style pack (use this one) 0.5 MB
voices/af_heart.bin legacy style pack 0.5 MB

Interface

Same signature as onnx-community/Kokoro-82M-v1.0-ONNX:

name type shape
input_ids (in) int64 [1, T] phoneme ids from config.json vocab, wrapped in 0 … 0, T ≤ 512
style (in) float32 [1, 256] = voice[num_phonemes - 1]
speed (in) float32 [1], 1.0 = normal
waveform (out) float32 [N] 24 kHz audio

Voice files are raw float32 arrays of shape [510, 1, 256].

Usage

pip install onnxruntime "misaki[en]" soundfile huggingface_hub
python -m spacy download en_core_web_sm
from inference import Kokoro7M   # inference.py in this repo
import soundfile as sf

tts = Kokoro7M()                 # or Kokoro7M(model="onnx/model_fp16.onnx")
sf.write("out.wav", tts("Hello, this is a small English voice."), 24000)

Use af_msa.bin: that is the style pack the student was distilled against. af_heart.bin degrades quality and is included only for compatibility.

Validation

Six sentences of 35 to 119 phonemes, ONNX compared against the PyTorch model:

  • Predicted durations, and therefore output lengths, match PyTorch exactly for fp32 and fp16.
  • The log-spectrogram L1 between ONNX fp32 and PyTorch is 0.33 to 0.39. Two PyTorch runs differ by 0.31 to 0.38, because the decoder's source excitation is random. fp16 adds roughly 0.01 on top of that.
  • Whisper-base.en transcribes ONNX fp32 and fp16 the same as PyTorch. The only difference is the made-up name "Kokoro", which Whisper hears as "Picoro" or "Kakoro" in every runtime, PyTorch included.
  • CPU real-time factor of the fp32 graph with 4 threads (onnxruntime 1.30, Apple Silicon) is about 0.027, or roughly 37x faster than realtime.

Listen: samples/. export_onnx.py reproduces the export (torch 2.5.1, opset 20). It expects the original repo downloaded to ./src.

Downloads last month
1
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vlapky/Kokoro-7M-Distill-ONNX

Quantized
(4)
this model