Instructions to use delebash/chatterbox-nano-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Chatterbox
How to use delebash/chatterbox-nano-GGUF with Chatterbox:
# pip install chatterbox-tts import torchaudio as ta from chatterbox.tts import ChatterboxTTS model = ChatterboxTTS.from_pretrained(device="cuda") text = "Ezreal and Jinx teamed up with Ahri, Yasuo, and Teemo to take down the enemy's Nexus in an epic late-game pentakill." wav = model.generate(text) ta.save("test-1.wav", wav, model.sr) # If you want to synthesize with a different voice, specify the audio prompt AUDIO_PROMPT_PATH="YOUR_FILE.wav" wav = model.generate(text, audio_prompt_path=AUDIO_PROMPT_PATH) ta.save("test-2.wav", wav, model.sr) - Notebooks
- Google Colab
- Kaggle
Chatterbox Nano β GGUF for audio.cpp, with voice cloning
8-bit and 16-bit GGUF conversions of ResembleAI/chatterbox-nano
(revision 71ccd1d0081b430592cea481f4307e764e07bc64) for audio.cpp.
Turbo's architecture with a GPT-2 small T3 (110M, 12 heads against Turbo's 16), the same tokenizer and 19 inline tags, and the same meanflow S3Gen decoder. English.
This is an unofficial conversion. The model, its weights and its licence are Resemble AI's. audio.cpp's own model repo has no Nano package. These files keep the voice encoder, the S3 speech tokenizer and the CAMPPlus speaker encoder, so the model clones a voice from a reference clip as well as speaking its built-in voice.
The files
| File | Size | SHA-256 |
|---|---|---|
chatterbox-nano-q8_0.gguf |
616,250,532 bytes | cb320915e43cf318a052392ed742504169886a3bcbf2b4871c0679aca73aee31 |
chatterbox-nano-f16.gguf |
894,652,290 bytes | a0f68a2dad58c115b38ce87c1ec32c15146db8668062423f5c23f2f258808134 |
Standalone GGUFs: the audio.cpp model spec and the tokenizer files are embedded.
Which audio.cpp runs them
Voice cloning, and Nano's 12 attention heads, need JustVoice's copy of audio.cpp,
delebash/audio.cpp, from release v0.9.0-jv.3 on.
Usage
audiocpp_cli --task tts --family chatterbox_turbo --model chatterbox-nano-q8_0.gguf --backend cuda \
--voice-ref speaker.wav --text "The harbour lights came on one by one." --out out.wav
Leave out --voice-ref for the built-in voice. The reference clip must be longer than 5
seconds. As upstream's tts_turbo.py does, it is loudness-normalised to β27 LUFS, the first
15 s give the T3's 375-token prompt and the first 10 s the decoder's.
How it was made
With convert_chatterbox_turbo.py, from Resemble's own files
(t3_nano_v1.safetensors, s3gen_meanflow.safetensors, ve.safetensors, conds.pt and the tokenizer files),
and audiocpp_gguf built from the same commit:
python3 tools/community_models/chatterbox_turbo/convert_chatterbox_turbo.py \
--checkpoint <snapshot> --output chatterbox-nano-q8_0.gguf --type q8_0
The 16-bit file is --type f16 (audio.cpp ships core Chatterbox at f16). No PyTorch is used:
the built-in voice is read from conds.pt directly. The three encoders are byte-identical to
core Chatterbox's own (ResembleAI/chatterbox's ve.safetensors and s3gen.safetensors), so
audio.cpp's core Chatterbox code computes a voice from them. The rest is renamed to what
audio.cpp's Turbo loaders read: GPT-2's Conv1D weights transposed, the vocoder's weight norm
folded, the head count stored as t3/hparams.num_heads.
What was checked
On a CPU build of that commit, one English line, seed 7, q8_0. The references were two
synthetic voices (Kokoro's af_heart, a US woman, and bm_george, a UK man; 9.5 s and 10.3 s).
"Similarity" is the cosine of Resemble's voice-encoder embedding against each reference: own /
other. Qwen3-ASR 1.7B read every render back.
| Render | Similarity | Median pitch | Read-back |
|---|---|---|---|
| Built-in voice | β | 222 Hz | exact |
| Clone of the US woman | 0.939 / 0.590 | 200 Hz (reference 202) | exact |
| Clone of the UK man | 0.929 / 0.582 | 125 Hz (reference 142) | exact |
Clones 0.946 and 0.918 at f16. A clip of 3 seconds is refused by name. Not checked: long-form output, listening
tests, real human voices, GPU speed.
Licence
MIT, the original's licence; the full text is in LICENSE. Changes from the original: the
weights were converted to GGUF β quantised to q8_0 or stored as f16 β with the tensors
renamed and the vocoder's weight norm folded as described above, and the tokenizer files and the
built-in voice embedded in the GGUF.
- Downloads last month
- 134
8-bit
16-bit
Model tree for delebash/chatterbox-nano-GGUF
Base model
ResembleAI/chatterbox-nano