Chatterbox (English)

Resemble AI's Chatterbox TTS (English), exported for loom.cpp: a Llama token model and a flow-matching decoder in one file. Encodes text itself and has a built-in voice.

This is a loom.cpp export: a single self-describing GGUF that carries its own graph topologies, tokenizer (if any) and driver script, produced by loom-exporter.

Original model

Exported from ResembleAI/chatterbox. Weights are unmodified; this repo packages the same parameters into loom.cpp's GGUF format.

License

mit, inherited from the base model above.

Language(s)

en

Usage

Run it with loom-py -- loom-py-rt on PyPI:

pip install -U "loom-py-rt[hub]"
import loom

model = loom.Model.from_pretrained("loom-ai-org/chatterbox-loom")

# This model encodes text itself -- no phonemiser needed at all.
print(model.tokenizer)                       # kind, vocabulary size, default language

# sample_rate=24000: a rate is not something a checkpoint necessarily carries, so it is a value
# you have to know from the model's documentation and pass. It is used only if the GGUF declares none;
# a wrong rate does not fail, it plays the voice at the wrong speed.
audio = model.text2speech.infer("hello world", sample_rate=24000)
audio.save("out.wav")

# That uses the voice the file itself defaults to. Whether it carries others is under "Known
# limitations" (and, where it does, a section below says how to pick one).

The layer underneath

The call above is the high-level door: one per task, named for the modality pair it maps between, with the windowing, sampling and assembly this model needs already applied. Under it, model.infer(...) passes your arguments straight to the driver this GGUF embeds -- which is where you go for a knob the door does not name.

model.driver_source prints that driver, including a header comment documenting every argument it accepts for this model, and is the authority on it. See loom-py for the API and loom.cpp for what the engine does between the two.

Known limitations

No watermark. Resemble AI's own package runs every output through Perth, an imperceptible neural watermark that lets synthetic speech be detected later. This export does NOT include it, deliberately: Perth is a separate neural model applied after synthesis, and this build targets local inference and small devices, where a smaller, faster model is the point. Audio from this file therefore carries no watermark and cannot be identified as synthetic by Perth's detector. If you distribute generated speech, disclose that it is synthetic yourself.

One voice, built in. Synthesis uses the checkpoint's own default voice (conds.pt). Cloning a new voice needs the voice encoder, the S3 speech tokenizer and CAMPPlus, which this export does not carry.

Sampled by default. Like the reference, the speech-token model samples (temperature 0.8, min_p 0.05, repetition penalty 1.2, classifier-free guidance 0.5), so two calls differ; pass seed to reproduce one, or temperature=0 for the deterministic guided decode the export is verified with (waveform within 2.5e-05 of the reference).

English only. The multilingual and Turbo checkpoints are different models and are not this file. Event tags such as [laughter] and [sigh] in the text are passed to the model as the reference passes them.

Files

  • chatterbox.gguf -- the model, exported with loom-exporter.
Downloads last month
29
GGUF
Model size
0.7B params
Architecture
loom-chatterbox
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for loom-ai-org/chatterbox-loom

Quantized
(33)
this model

Collection including loom-ai-org/chatterbox-loom