Qwen3.8-27B-oQ6e-mtp

6-bit MLX quantization of Qwen/Qwen3.8-27B (hybrid Gated DeltaNet + Gated Attention, native vision, 262K context, thinking mode with preserved reasoning), sized for 64 GB-class Apple Silicon. Made with oQ (oMLX v0.6.0.dev1) mixed-precision quantization.

Sibling repos:

Quantization details

  • Model type: qwen3_5
  • Bits: 6 (effective 6.8 bpw mixed precision), group size 64 — ~22 GiB weights, leaving KV-cache headroom under a 64 GB machine's default GPU-memory limit
  • Enhanced quantization (oQe): imatrix-calibrated (1024 samples) — affine quantization weighted by activation importance
  • MTP weights preserved (mtp.* tensors + config) — multi-token-prediction / Lightning MTP works after quantization
  • Non-quant weight dtype: bfloat16 (matches the base model; the safe choice on M3/M4/M5)
  • Vision components included (not a text-only strip)
  • Format: MLX safetensors

Recommended sampling (per the Qwen3.8 model card)

Mode temperature top_p
Thinking (default) 1.0 0.95
Instruct (non-thinking) 0.7 0.80

Thinking controls via chat_template_kwargs:

  • enable_thinking (default true)
  • preserve_thinking (default true) — keeps reasoning traces across multi-turn history
  • reasoning_effort: xhigh (default) / medium / low — in our testing (on the 8-bit sibling), medium reduced thinking volume ~25% with no loss on agentic tasks

Tested

The 8-bit sibling of this quant was validated 2026-08-14 on an M2 Mac Studio: 0 stalled turns in multi-turn agentic tool use with thinking ON and thinking blocks fed back into history (synthetic battery + a 12-turn live session at 27 messages of history) — see the sibling repo's card for details. This 6-bit build uses the same imatrix calibration; it has not yet been independently run through the same battery.

Downloads last month
428
Safetensors
Model size
7B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for evsinlb/Qwen3.8-27B-oQ6e-mtp

Base model

Qwen/Qwen3.8-27B
Quantized
(678)
this model