Chatterbox Nano β€” GGUF for audio.cpp, with voice cloning

8-bit and 16-bit GGUF conversions of ResembleAI/chatterbox-nano (revision 71ccd1d0081b430592cea481f4307e764e07bc64) for audio.cpp. Turbo's architecture with a GPT-2 small T3 (110M, 12 heads against Turbo's 16), the same tokenizer and 19 inline tags, and the same meanflow S3Gen decoder. English.

This is an unofficial conversion. The model, its weights and its licence are Resemble AI's. audio.cpp's own model repo has no Nano package. These files keep the voice encoder, the S3 speech tokenizer and the CAMPPlus speaker encoder, so the model clones a voice from a reference clip as well as speaking its built-in voice.

The files

File Size SHA-256
chatterbox-nano-q8_0.gguf 616,250,532 bytes cb320915e43cf318a052392ed742504169886a3bcbf2b4871c0679aca73aee31
chatterbox-nano-f16.gguf 894,652,290 bytes a0f68a2dad58c115b38ce87c1ec32c15146db8668062423f5c23f2f258808134

Standalone GGUFs: the audio.cpp model spec and the tokenizer files are embedded.

Which audio.cpp runs them

Voice cloning, and Nano's 12 attention heads, need JustVoice's copy of audio.cpp, delebash/audio.cpp, from release v0.9.0-jv.3 on.

Usage

audiocpp_cli --task tts --family chatterbox_turbo --model chatterbox-nano-q8_0.gguf --backend cuda \
  --voice-ref speaker.wav --text "The harbour lights came on one by one." --out out.wav

Leave out --voice-ref for the built-in voice. The reference clip must be longer than 5 seconds. As upstream's tts_turbo.py does, it is loudness-normalised to βˆ’27 LUFS, the first 15 s give the T3's 375-token prompt and the first 10 s the decoder's.

How it was made

With convert_chatterbox_turbo.py, from Resemble's own files (t3_nano_v1.safetensors, s3gen_meanflow.safetensors, ve.safetensors, conds.pt and the tokenizer files), and audiocpp_gguf built from the same commit:

python3 tools/community_models/chatterbox_turbo/convert_chatterbox_turbo.py \
  --checkpoint <snapshot> --output chatterbox-nano-q8_0.gguf --type q8_0

The 16-bit file is --type f16 (audio.cpp ships core Chatterbox at f16). No PyTorch is used: the built-in voice is read from conds.pt directly. The three encoders are byte-identical to core Chatterbox's own (ResembleAI/chatterbox's ve.safetensors and s3gen.safetensors), so audio.cpp's core Chatterbox code computes a voice from them. The rest is renamed to what audio.cpp's Turbo loaders read: GPT-2's Conv1D weights transposed, the vocoder's weight norm folded, the head count stored as t3/hparams.num_heads.

What was checked

On a CPU build of that commit, one English line, seed 7, q8_0. The references were two synthetic voices (Kokoro's af_heart, a US woman, and bm_george, a UK man; 9.5 s and 10.3 s). "Similarity" is the cosine of Resemble's voice-encoder embedding against each reference: own / other. Qwen3-ASR 1.7B read every render back.

Render Similarity Median pitch Read-back
Built-in voice β€” 222 Hz exact
Clone of the US woman 0.939 / 0.590 200 Hz (reference 202) exact
Clone of the UK man 0.929 / 0.582 125 Hz (reference 142) exact

Clones 0.946 and 0.918 at f16. A clip of 3 seconds is refused by name. Not checked: long-form output, listening tests, real human voices, GPU speed.

Licence

MIT, the original's licence; the full text is in LICENSE. Changes from the original: the weights were converted to GGUF β€” quantised to q8_0 or stored as f16 β€” with the tensors renamed and the vocoder's weight norm folded as described above, and the tokenizer files and the built-in voice embedded in the GGUF.

Downloads last month
134
GGUF
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for delebash/chatterbox-nano-GGUF

Quantized
(10)
this model