agnosticeng/Qwen3.8-27B-4bit

4-bit MLX quantization of Qwen/Qwen3.8-27B, with the Qwen3.8 MTP head bundled at mtp/ for speculative decoding.

Quantization

  • Affine 4-bit, group size 64 (~4.7 bits/weight, ~15 GB)
  • Converted with mlx-vlm; vision tower included (image-text-to-text)

MTP head

mtp/weights.safetensors contains the Qwen3.8-27B multi-token-prediction sidecar (4-bit, mtp.-prefixed keys).

Downloads last month
-
Safetensors
Model size
27B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for agnosticeng/Qwen3.8-27B-4bit

Base model

Qwen/Qwen3.8-27B
Quantized
(994)
this model