Instructions to use OpenVoiceOS/phoonnx-pocket-tts with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Pocket-TTS
How to use OpenVoiceOS/phoonnx-pocket-tts with Pocket-TTS:
from pocket_tts import TTSModel import scipy.io.wavfile tts_model = TTSModel.load_model("OpenVoiceOS/phoonnx-pocket-tts") voice_state = tts_model.get_state_for_audio_prompt( "hf://kyutai/tts-voices/alba-mackenna/casual.wav" ) audio = tts_model.generate_audio(voice_state, "Hello world, this is a test.") # Audio is a 1D torch tensor containing PCM data. scipy.io.wavfile.write("output.wav", tts_model.sample_rate, audio.numpy()) - Notebooks
- Google Colab
- Kaggle
Pocket TTS β ONNX mirror for phoonnx
ONNX weights for Kyutai Pocket TTS, packaged for the
phoonnx pockettts engine.
Pocket TTS is a 100-million-parameter text-to-speech model. It runs faster than real time on two CPU cores. This mirror holds one weight bundle per language, in full precision and in 8-bit, together with the published speaker states.
Languages
| Directory | Language | Transformer layers |
|---|---|---|
english_2026-04 |
English | 6 |
french_24l |
French | 24 |
german |
German | 6 |
italian |
Italian | 6 |
portuguese |
Portuguese | 6 |
spanish |
Spanish | 6 |
Contents of a bundle
| File | Purpose |
|---|---|
bundle.json |
Sample rate, latent size, chunk limit and the two state manifests |
tokenizer.model |
SentencePiece tokenizer, 4000 units |
bos_before_voice.npy |
Embedding prepended to a cloned voice |
text_conditioner.onnx |
Token ids to text embeddings |
flow_lm_main.onnx |
Autoregressive transformer with explicit state |
flow_lm_flow.onnx |
Flow network for one latent frame |
mimi_decoder.onnx |
Latent frames to 24 kHz audio |
mimi_encoder.onnx |
Reference audio to latent frames, for voice cloning |
voices/<name>.safetensors |
Published speaker states, 26 per language |
Every graph except the encoders also ships as *_int8.onnx.
Graph interfaces
text_conditioner : token_ids -> embeddings
flow_lm_main : sequence, text_embeddings, state_* -> conditioning, eos_logit, out_state_*
flow_lm_flow : c, s, t, x -> flow_dir
mimi_decoder : latent, state_* -> audio_frame, out_state_*
mimi_encoder : audio -> latents
Audio is 24 kHz mono. The model produces 12.5 latent frames per second, and each frame
decodes to 1920 samples. The state_* tensors are described by the state manifests in
bundle.json: each entry gives the tensor name, shape, dtype and how to fill it at the
start of a stream.
A voice is a state, not an embedding vector: it is the transformer state after the model
has consumed a speaker's audio. The manifests map a voices/*.safetensors file onto the
state tensors of flow_lm_main.
Usage
from phoonnx.model_manager import TTSModelManager
manager = TTSModelManager()
voice = manager.download_voice_by_id("pockettts/en/alba").load()
audio = voice.synthesize_wav("Hello world.")
Attribution
The model is the work of Kyutai, released as
kyutai/pocket-tts. The ONNX graphs were
exported by the community and are mirrored from
KevinAHM/pocket-tts-onnx. The speaker
states come from
kyutai/pocket-tts-without-voice-cloning.
This repository adds no new weights. It exists so that phoonnx can fetch a stable set of files from one place.
License
The weights keep their upstream license, Creative Commons Attribution 4.0 International
(CC BY 4.0). The full text is in LICENSE. Give credit to Kyutai when you use them.
Kyutai asks that the model is not used for voice impersonation without consent, for deception, or for unlawful or abusive content.
Model tree for OpenVoiceOS/phoonnx-pocket-tts
Base model
kyutai/pocket-tts