Instructions to use PowerBeef02/Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use PowerBeef02/Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bit PowerBeef02/Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
PowerBeef02/Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bit
Vocello production artifact. Derived from
mlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bit
at revision 5c390979e4b93af5f2932f90742ca99c7dd04687 (itself an MLX conversion of the
corresponding Qwen/Qwen3-TTS checkpoint), with one change:
the 622 MB BF16 talker.model.text_embedding tensor is quantized to affine 8-bit
(group size 64) by Vocello's pinned conversion tooling
(scripts/convert_text_embedding_8bit.py, python-mlx 0.32.0). Every other tensor is
byte-identical to the source revision.
Measured on the Vocello support floor (Mac mini M2 8 GB): about 278 MB less resident
memory and 292 MB less download per Speed artifact at parity real-time factor, with clean
deterministic audio QC. Measurement details live in the
Vocello repository (benchmarks/OPTIMIZATION.md §N).
These files are consumed by Vocello's fail-closed model catalog, which pins this repository at an exact revision with per-file SHA-256 digests.
- Downloads last month
- 45
4-bit