Qwen 3.8 27B MTPLX 5-bit

Every weight matrix at 5-bit, sensitive parts at 16-bit, native multi-token-prediction head kept, so MTPLX still drafts ahead and verifies in one pass.

Speeds

Measured on an Macbook M5 (NOT pro or max) with 32GB unified memory, fans verified at max, single stream, generation running to the model's own stop, official Qwen 3.8 sampling (temperature 1.0, top-p 0.95, top-k 20).

How it is built

  • Every weight matrix at 5-bit with 64-weight groups.

  • The GDN convolution kernels and recurrent state parameters, every norm, and the whole MTP head stay 16-bit.

  • Download 19.4 GB

  • Context window 262,144 tokens

  • MTP depth 3

  • Sampling: temperature 1.0, top-p 0.95, top-k 20 (the official Qwen 3.8 contract)

Use it

You want 32 GB of unified memory or more for this one.

Command line:

pip install mtplx

mtplx serve --model FlatFootInternational/qwen3.8-27b-MTPLX-5bit
Downloads last month
292
Safetensors
Model size
5B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FlatFootInternational/qwen3.8-27b-MTPLX-5bit

Base model

Qwen/Qwen3.8-27B
Quantized
(935)
this model