Breeze-TTS-2 GGUF

GGUF conversions of BreezeBlue/Breeze-TTS-2 for Breeze-TTS-2.cpp, a C++ reimplementation running on ggml with a Vulkan backend, so it works on NVIDIA, AMD and Intel GPUs and falls back to CPU.

Bilingual English and Mandarin, 24 kHz, around 1.2x realtime at Q8_0 on an RTX 3060.

These files will not load in llama.cpp. They need the Breeze-TTS-2.cpp runtime, which implements all four stages of the model: the T5Gemma2 text encoder, the Qwen3 backbone, the 15 step depth decoder and the vocoder.

Files

File Size Notes
breeze-tts-2-f16.gguf 5.9 GB Reference quality, unquantized
breeze-tts-2-q8_0.gguf 3.3 GB Recommended. No audible loss against F16
breeze-tts-2-q6_k.gguf 2.9 GB
breeze-tts-2-q4_k.gguf 2.4 GB Smallest safe choice, holds up well
breeze-tts-2-q8_0-dd4.gguf 3.2 GB Experimental, Q8_0 base with a Q4_K depth decoder
breeze-tts-2-q8_0-dd2.gguf 3.1 GB Experimental, Q8_0 base with a Q2_K depth decoder
breeze-tts-2-q4_k-dd2.gguf 2.3 GB Experimental, Q4_K base with a Q2_K depth decoder

Approximate VRAM is about 1 GB above the file size.

About the -dd variants

Everything except the -dd files keeps the depth decoder at higher precision than the rest of the model. The -dd variants quantize it too, which is why they are smaller.

The depth decoder runs 15 sequential steps for every single frame of audio, so on hardware where that is the bottleneck rather than memory bandwidth, shrinking it can speed generation up noticeably. That is the reason these exist and it is worth benchmarking on your own card.

The tradeoff is that depth codes feed back into the backbone every frame, so quantization error compounds as generation continues. Output holds up early and then drifts progressively muffled and thin past roughly 45 seconds of continuous generation. Short lines and dialogue are fine. Long narration is not, and the failure is gradual rather than obvious, so it is easy to miss on quick tests.

Treat them as experimental. If in doubt, use q8_0 or q4_k.

Usage

git clone --recursive https://github.com/HoppouAI/Breeze-TTS-2.cpp
cd Breeze-TTS-2.cpp
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j
# invent a voice from a description
build/breeze-cli breeze-tts-2-q8_0.gguf \
  --text "Welcome aboard. Your journey begins now." \
  --instruction "A warm, thoughtful young woman with a clear, calm delivery." \
  --output design.wav

# clone a voice from a clip
build/breeze-cli breeze-tts-2-q8_0.gguf \
  --text "It is good to hear your voice again." \
  --ref-audio ref_voice.wav --ref-text "The harbour lights came on one by one as the evening tide began to turn." \
  --output clone.wav

Or run the server, which has a web UI built in plus HTTP and WebSocket streaming:

build/breeze-server breeze-tts-2-q8_0.gguf --host 127.0.0.1 --port 8080 --webui

ref_voice.wav in this repo is a sample reference clip. Its transcript is "The harbour lights came on one by one as the evening tide began to turn."

Vocal events

Inline tags in round brackets produce non speech sounds: (laugh), (sigh), (cough), (clears throat), and [笑] or [叹气] in Chinese. The vocabulary is free form rather than a fixed token list, so descriptive tags like (nervous chuckle) often work.

They usually need --cfg-scale 2 to 3 to actually fire. At the default of 1.0 the model treats a tag as a suggestion and tends to read straight past anything outside the common set.

Voice conversion

The runtime can also respeak an existing recording in a different voice, keeping the original timing, phrasing and emphasis while changing only the speaker. This is not part of the upstream model, it falls out of how the codec separates semantic content from acoustic detail.

It is experimental. Pitch is regenerated rather than copied, so a converted vocal is re-sung in the target voice's own register instead of at the source's. Whether a tune survives varies clip to clip, and keep_acoustic 1 or 2 copies the lowest acoustic codebooks from the source to pull more of the original contour through. Judge it by ear, and leave keep_acoustic at 0 for ordinary speech.

Conversion

Produced with scripts/convert_hf_to_gguf.py and breeze-quantize from the repo. The source download must include the audio_tokenizer/ directory, which holds the vocoder that the model actually uses at inference time.

License

Weights are governed by the BreezeBlue Research and Non-Commercial License from the original model. Converting to GGUF does not change that. The Breeze-TTS-2.cpp source code is Apache 2.0.

You are responsible for complying with the weight license and for obtaining consent for any reference audio or voices you use.

Downloads last month
-
GGUF
Model size
3B params
Architecture
breeze-tts-2
Hardware compatibility
Log In to add your hardware

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for HoppouAI/Breeze-TTS-2.cpp

Quantized
(5)
this model