lukaskremla/Qwen3.8-27B-MTP-3bit-MLX

This repository contains a 3-bit affine quantization of the native Multi-Token Prediction (MTP) tensors extracted from Qwen/Qwen3.8-27B and converted into the standalone MLX drafter format expected by mlx-vlm.

This is not a standalone language model. It must be loaded as the draft model alongside a compatible Qwen3.8-27B target model.

This 3-bit drafter can be used with the 2-bit, 3-bit, 4-bit, 5-bit, 6-bit and 8-bit MLX target models. Quantizations don't need to match.

You can download the vision capable or text-only Qwen 3.8 27B weights in a variety of MLX quantizations from this collection Qwen 3.8 27B MLX-Quants (Vision, Text-Only & MTP).

Hugging Face might render incorrect parameter counts for this model, it is a common display bug for MLX quants.

Credits

Converted to MLX format from Qwen/Qwen3.8-27B using mlx-vlm version 0.6.13.

Downloads last month
128
Safetensors
Model size
53.1M params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including lukaskremla/Qwen3.8-27B-MTP-3bit-MLX