MTPLX.COM: 2 to 3x speedup. The fastest way to run models on a Mac.

Qwen 3.8 27B Bare Speed

Quickest burst chat speeds. Lower quality and slower on long coding tasks.

The fastest of the three MTPLX Qwen 3.8 builds and the smallest download. Qwen3.8-27B in flat 4-bit with its native multi-token-prediction head kept, so MTPLX drafts ahead and verifies in one pass. If you want the snappiest chat on a Mac and can live with a rougher quant, this is it. For coding, pick Optimized Speed.

Speeds

Measured on an M5 Max, fans verified at max, single stream, generation running to the model's own stop, official Qwen 3.8 sampling (temperature 1.0, top-p 0.95, top-k 20).

Run tok/s
Coding task, medium reasoning, mtplx serve 65.2
Same task inside the MTPLX Mac app 64.4
Long reasoning at xhigh, 34k and 37k token answers 35.7 and 32.0
One 52,740-token answer, 27.2 minutes, ended at the model's own stop 32.4 sustained

Same night, same task: the previous MTPLX flagship Qwen 3.6 27B Optimized Speed V2 ran 59.9 to 60.1 tok/s. oMLX 0.5.7 serving its own Qwen 3.8 4-bit MTP quant ran 63.3. LM Studio on the 52k-token long answer ran 17.40 tok/s against 32.4 here.

Draft acceptance on the coding task by depth: 0.95, 0.86, 0.78. Verify cost 44 ms per round.

How it is built

  • Every weight matrix at 4-bit with 64-weight groups. Nothing promoted.
  • The GDN convolution kernels and recurrent state parameters, every norm, and the whole MTP head stay 16-bit.
  • KL divergence to the original bf16 model on our coding battery: 0.0376. Optimized Speed is 1.7x closer, Optimized Quality 36x closer. That is the trade you make for the speed.
Download 16.0 GB
Peak unified memory (measured, this artifact) 17.0 GB
Context window 262,144 tokens
MTP depth 3
Sampling temperature 1.0, top-p 0.95, top-k 20 (the official Qwen 3.8 contract)
Draft sampler temperature 0.6 (measured winner for this build, 46.1 vs 42.4 tok/s)

The tuned depth and draft settings ship inside mtplx_runtime.json. MTPLX reads them on load. The draft sampler is a speed knob only: MTPLX accepts drafts with the probability-ratio rule plus residual resampling, so the output follows the model's own distribution at any temperature. Reasoning effort levels (xhigh, medium, low) work, and preserved thinking flows through the MTP path.

Use it

Mac app: download at mtplx.com, pick "Qwen 3.8 27B Bare Speed".

Command line:

pip install mtplx
mtplx serve --model Youssofal/Qwen3.8-27B-MTPLX-Bare-Speed

Siblings: Optimized Speed (recommended for coding) and Optimized Quality (8-bit, perfect quality). On an M1 or M2 Mac use the FP16 build of this model.

Downloads last month
6
Safetensors
Model size
4B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Youssofal/Qwen3.8-27B-MTPLX-Bare-Speed

Base model

Qwen/Qwen3.8-27B
Quantized
(372)
this model