Pocket TTS for Glade

Model assets and runtime configuration for Kyutai Pocket TTS in Glade. Ordinary text produces mono 24 kHz PCM, with complete and incremental synthesis interfaces. These are the standard six-layer checkpoints.

Languages and voices

Languages are independent downloads. languages.json declares exact file groups, shared files and sizes; metadata.json declares the profile inventory. Each profile includes its original language weights, Mimi codec, tokenizer, text preparation rules and official preset conditioning. English is the default.

Language Official presets Download
English Alba, Anna 121 MB
French Estelle 118 MB
German Juergen 117 MB
Portuguese Rafael 117 MB
Italian Giovanni 116 MB
Spanish Lola 117 MB
Dutch Daan 116 MB

All languages together total approximately 823 MB. Adding a language does not require downloading the other profiles. Glade uses SwiftModelHub to verify files against Hugging Face's native checksums and reuse unchanged files during updates.

Execution

Each profile contains two source .aimodel assets: W8A16 language inference and an FP16 Mimi audio decoder. Both prefer ANE. The language graph retains six causal Transformer layers and the original one-step LSD flow head, producing an 80 ms latent at each autoregressive step. The full-history language cache has 512 positions, including the selected preset's prefix; prefill uses 32 positions.

The codec supports 1/2/4/8 consecutive latents with shared weights and preserves its native 250-position causal attention context and convolution state. Independent utterances are processed sequentially. Original tokenization, character replacements, sentence splitting, 50-token target, EOS conventions and short fade-in are retained. Unsupported excess capacity reports an error. No compiler specialization caches or reference recordings are included.

Measured performance

Complete reference text of JFK's “We choose to go to the Moon” speech: 2,201 words, one submitted text request, all 90 native chunks, English/Alba, temperature 0.3, seed 42 and resident execution. Release common-session medians of three warmed passes include text planning, tokenization, all model calls and host state work; preparation, warmup and WAV writing are separate.

Device Generated audio Text-to-audio RTFx
M3 MacBook Air, 16 GB, macOS 27.0.1 10 min 42 s 30.31 s 21.2×
iPhone 15 Pro Max, A17 Pro, 8 GB, iOS 27.0.1 10 min 42 s 40.71 s 15.8×

RTFx is generated audio seconds divided by generation seconds. Phone runs stayed at nominal thermal state and produced identical PCM across repetitions. Its sampled client peak was 409 MiB, including retained output PCM; compiler and external service memory is excluded. Seeds do not imply identical waveforms between devices or the original PyTorch sampler.

Native-language short samples use the same text, seed and preset on both devices:

Language · preset M3 RTFx iPhone RTFx iPhone client peak
en · Alba 23.8× 12.8× 286 MiB
en · Anna 22.9× 13.0× 268 MiB
fr · Estelle 22.4× 11.1× 270 MiB
de · Juergen 25.3× 13.5× 284 MiB
pt · Rafael 24.6× 12.2× 269 MiB
it · Giovanni 23.8× 13.8× 273 MiB
es · Lola 23.1× 11.6× 276 MiB
nl · Daan 23.3× 12.8× 276 MiB

Preparation and qualification

Initial observed phone preparation for each language was 10.2–11.1 s, approximately half language graph and half codec. This is separate from generation; no cache invalidation was performed, so it is not a pristine-cache measurement. Anna reuses the English assets and took 0.08 s in a subsequent process. Full-speech cached loading measured 1.14 s on Mac and 0.20 s on phone.

Seven separate phone captures show one contained ANE prediction for every observed language/codec graph call and zero target GPU intervals. Short captures exercise codec batches 2/4/8; the full English control also exercises batch 1, whose placement was not separately traced in this candidate. Hardware intervals establish execution, not arithmetic occupancy or energy consumption.

Native text processing matches 70 original-source cases across all seven languages. FP32 graph rewrites match source controls below 3e-6 relative RMS. Post-quantization controls over 64 teacher-forced latents per preset, using each implementation's accumulated KV state, measure 1.7–2.4% aggregate latent RMS against source FP32. They do not establish free-running waveform equivalence. Cohere checks of sample outputs were coherent with occasional recognition/word substitutions. General perceptual and voice-similarity scores are not claimed. Watch qualification is not claimed.

Attribution

Weights: Kyutai, 3e82814a68665eec246ff649b14c71331f955c06, CC-BY-4.0. Tokenizers and official preset states: kyutai/pocket-tts-without-voice-cloning, 1e08e6a23401048648a9fdcfde2f89348215c2a7, CC-BY-4.0. Original implementation: Kyutai, 41cbc84af539ea78a804ffca5f9c6edc1a22ce44, MIT. Licenses and preset attribution are included. Conversion and quantization change inference representation, not authorship or ownership.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for coder543/pocket-tts-glade

Quantized
(53)
this model

Collection including coder543/pocket-tts-glade