PowerBeef02/Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bit

Vocello production artifact. Derived from mlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bit at revision 5c390979e4b93af5f2932f90742ca99c7dd04687 (itself an MLX conversion of the corresponding Qwen/Qwen3-TTS checkpoint), with one change: the 622 MB BF16 talker.model.text_embedding tensor is quantized to affine 8-bit (group size 64) by Vocello's pinned conversion tooling (scripts/convert_text_embedding_8bit.py, python-mlx 0.32.0). Every other tensor is byte-identical to the source revision.

Measured on the Vocello support floor (Mac mini M2 8 GB): about 278 MB less resident memory and 292 MB less download per Speed artifact at parity real-time factor, with clean deterministic audio QC. Measurement details live in the Vocello repository (benchmarks/OPTIMIZATION.md §N).

These files are consumed by Vocello's fail-closed model catalog, which pins this repository at an exact revision with per-file SHA-256 digests.

Downloads last month
45
Safetensors
Model size
0.4B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PowerBeef02/Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bit

Quantized
(1)
this model