Instructions to use coder543/pocket-tts-glade with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Pocket-TTS
How to use coder543/pocket-tts-glade with Pocket-TTS:
from pocket_tts import TTSModel import scipy.io.wavfile tts_model = TTSModel.load_model("coder543/pocket-tts-glade") voice_state = tts_model.get_state_for_audio_prompt( "hf://kyutai/tts-voices/alba-mackenna/casual.wav" ) audio = tts_model.generate_audio(voice_state, "Hello world, this is a test.") # Audio is a 1D torch tensor containing PCM data. scipy.io.wavfile.write("output.wav", tts_model.sample_rate, audio.numpy()) - Notebooks
- Google Colab
- Kaggle
Pocket TTS for Glade
Model assets and runtime configuration for Kyutai Pocket TTS in Glade. Ordinary text produces mono 24 kHz PCM, with complete and incremental synthesis interfaces. These are the standard six-layer checkpoints.
Languages and voices
Languages are independent downloads. languages.json declares exact file groups,
shared files and sizes; metadata.json declares the profile inventory. Each
profile includes its original language weights, Mimi codec, tokenizer, text
preparation rules and official preset conditioning. English is the default.
| Language | Official presets | Download |
|---|---|---|
| English | Alba, Anna | 121 MB |
| French | Estelle | 118 MB |
| German | Juergen | 117 MB |
| Portuguese | Rafael | 117 MB |
| Italian | Giovanni | 116 MB |
| Spanish | Lola | 117 MB |
| Dutch | Daan | 116 MB |
All languages together total approximately 823 MB. Adding a language does not require downloading the other profiles. Glade uses SwiftModelHub to verify files against Hugging Face's native checksums and reuse unchanged files during updates.
Execution
Each profile contains two source .aimodel assets: W8A16 language inference and
an FP16 Mimi audio decoder. Both prefer ANE. The language graph retains six causal
Transformer layers and the original one-step LSD flow head, producing an 80 ms
latent at each autoregressive step. The full-history language cache has 512
positions, including the selected preset's prefix; prefill uses 32 positions.
The codec supports 1/2/4/8 consecutive latents with shared weights and preserves its native 250-position causal attention context and convolution state. Independent utterances are processed sequentially. Original tokenization, character replacements, sentence splitting, 50-token target, EOS conventions and short fade-in are retained. Unsupported excess capacity reports an error. No compiler specialization caches or reference recordings are included.
Measured performance
Complete reference text of JFK's “We choose to go to the Moon” speech: 2,201 words, one submitted text request, all 90 native chunks, English/Alba, temperature 0.3, seed 42 and resident execution. Release common-session medians of three warmed passes include text planning, tokenization, all model calls and host state work; preparation, warmup and WAV writing are separate.
| Device | Generated audio | Text-to-audio | RTFx |
|---|---|---|---|
| M3 MacBook Air, 16 GB, macOS 27.0.1 | 10 min 42 s | 30.31 s | 21.2× |
| iPhone 15 Pro Max, A17 Pro, 8 GB, iOS 27.0.1 | 10 min 42 s | 40.71 s | 15.8× |
RTFx is generated audio seconds divided by generation seconds. Phone runs stayed at nominal thermal state and produced identical PCM across repetitions. Its sampled client peak was 409 MiB, including retained output PCM; compiler and external service memory is excluded. Seeds do not imply identical waveforms between devices or the original PyTorch sampler.
Native-language short samples use the same text, seed and preset on both devices:
| Language · preset | M3 RTFx | iPhone RTFx | iPhone client peak |
|---|---|---|---|
| en · Alba | 23.8× | 12.8× | 286 MiB |
| en · Anna | 22.9× | 13.0× | 268 MiB |
| fr · Estelle | 22.4× | 11.1× | 270 MiB |
| de · Juergen | 25.3× | 13.5× | 284 MiB |
| pt · Rafael | 24.6× | 12.2× | 269 MiB |
| it · Giovanni | 23.8× | 13.8× | 273 MiB |
| es · Lola | 23.1× | 11.6× | 276 MiB |
| nl · Daan | 23.3× | 12.8× | 276 MiB |
Preparation and qualification
Initial observed phone preparation for each language was 10.2–11.1 s, approximately half language graph and half codec. This is separate from generation; no cache invalidation was performed, so it is not a pristine-cache measurement. Anna reuses the English assets and took 0.08 s in a subsequent process. Full-speech cached loading measured 1.14 s on Mac and 0.20 s on phone.
Seven separate phone captures show one contained ANE prediction for every observed language/codec graph call and zero target GPU intervals. Short captures exercise codec batches 2/4/8; the full English control also exercises batch 1, whose placement was not separately traced in this candidate. Hardware intervals establish execution, not arithmetic occupancy or energy consumption.
Native text processing matches 70 original-source cases across all seven languages. FP32 graph rewrites match source controls below 3e-6 relative RMS. Post-quantization controls over 64 teacher-forced latents per preset, using each implementation's accumulated KV state, measure 1.7–2.4% aggregate latent RMS against source FP32. They do not establish free-running waveform equivalence. Cohere checks of sample outputs were coherent with occasional recognition/word substitutions. General perceptual and voice-similarity scores are not claimed. Watch qualification is not claimed.
Attribution
Weights: Kyutai, 3e82814a68665eec246ff649b14c71331f955c06, CC-BY-4.0.
Tokenizers and official preset states: kyutai/pocket-tts-without-voice-cloning,
1e08e6a23401048648a9fdcfde2f89348215c2a7, CC-BY-4.0. Original implementation:
Kyutai, 41cbc84af539ea78a804ffca5f9c6edc1a22ce44, MIT. Licenses and preset
attribution are included. Conversion and quantization change inference
representation, not authorship or ownership.
- Downloads last month
- -
Model tree for coder543/pocket-tts-glade
Base model
kyutai/pocket-tts