Qwen3.5-2B — 4-bit MLX, text-only layout

The text decoder of Qwen/Qwen3.5-2B (Apache-2.0), quantized to 4-bit with mlx-lm (group 64, affine) and re-keyed to the qwen3_5_text layout (model.* tree) so mlx-swift-lm's Qwen35TextModel + PEFT adapter loader consume it directly. The chat template is pinned to non-thinking rendering. Base weights only — no fine-tuning of any kind.

Downloads last month
94
Safetensors
Model size
0.3B params
Tensor type
BF16
·
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for augible/qwen3.5-2b-4bit-text-mlx

Finetuned
Qwen/Qwen3.5-2B
Quantized
(150)
this model