Paradee-8M for Glade
Model assets and runtime configuration for Paradee-8M v1.0 in Glade. This single-voice American English distillation of Kokoro accepts ordinary text or explicit phonemes and produces mono 24 kHz PCM with its trained Heart voice.
Text and execution
The approximately 22.9 MB download includes the American English Misaki frontend: trained spaCy tokenization/POS, contextual lexical processing, number handling and the original BART pronunciation fallback. No separate frontend download or Python runtime is needed.
INT8 text, prosody and acoustic weights use FP16 activations on ANE. The GPU waveform generator retains FP16 weights. CPU parameters use FP16 storage and FP32 recurrence/signal processing. The original 33-frame phase-locking filter remains enabled; utterance normalization excludes padding.
metadata.json specifies the 96-token/192-acoustic-frame capacities, masking,
compute preferences and phase filter. Four .aimodel assets and cpu.f16/cpu.json
supply the neural/audio stages. Text is phonemized in context and partitioned at
the model's capacity; duration overflow is retried with smaller parts before PCM
delivery. The complete input is retained. No compiled cache is supplied.
Measured performance
Complete 2,201-word reference text of JFK's “We choose to go to the Moon” speech, Heart voice, speed 1, seed 42, one synthesis lane, resident execution. Warmed Release text-to-audio medians of three passes include phonemization, planning, neural calls, recurrence and signal processing. Initial preparation, warmup and WAV writing are separate.
| Device | Generated audio | Text-to-audio | RTFx |
|---|---|---|---|
| M3 MacBook Air, 16 GB, macOS 27.0.1 | 13 min 30 s | 14.87 s | 54.4× |
| iPhone 15 Pro Max, A17 Pro, 8 GB, iOS 27.0.1 | 13 min 30 s | 9.59 s | 84.4× |
RTFx uses generated audio duration. The phone remained at nominal thermal state; sampled client peak was 266 MiB, including benchmark PCM and excluding compiler/service memory. All three phone outputs were byte-identical. A separate trace confirms ANE text/prosody/acoustic execution and GPU waveform generation.
Observed cached preparation was 0.30 s on Mac and 0.31 s on phone. Compiler caches were not purged; these are not cold specialization measurements. First use requires device specialization. Original frontend/numerical controls and sample recognition checks qualify functionality, not general perceptual scores. Watch qualification is not claimed.
Attribution
Weights: f662642d44c03c17588e4176469c54d462c0b623.
Original implementation: 9c8b4d7504cbee7e64de2d0341bb690f0b0ab708.
The original model/code and Misaki use Apache-2.0; trained frontend components
retain their bundled notices. No additional training is performed. Download
weights and configuration from one repository snapshot.
Model tree for coder543/paradee-8m-glade
Base model
yl4579/StyleTTS2-LJSpeech