Qwen3.6-27B-Opus-Distill-OptiQ-4bit

OptiQ mixed-precision MLX quant (~3.97 bpw, data-driven KL-sensitivity 4/8-bit) of TeichAI/Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2 (Apache-2.0) — a dense qwen3_5 (hybrid linear-attention) distilled on Claude Opus reasoning traces. Quantized with mlx_optiq.

Profile (64GB Apple-Silicon, 256K-agentic-coding study):

  • Clears 256K context: mx.get_peak_memory = 43.3 GB @ 256K (4-bit KV), retrieval 1.00.
  • Strong single-shot: HumanEval+ 100%, AIME 100%, math500 ~81%, LCB 80%, BFCL 0.94; aider-polyglot ~75%.
  • Dense-27B → slower decode than a sparse MoE at long context (~9 tok/s @ 256K).

Recommended sampling (op-temp from testing): temperature 0.3, top_p 0.95, top_k 20, min_p 0, presence_penalty 0, thinking enabled (thinking_budget 81920).

MLX/OptiQ quant + evaluation by @caslca. Base model + distillation by TeichAI (Apache-2.0).

Downloads last month
616
Safetensors
Model size
27B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for caslca/Qwen3.6-27B-Opus-Distill-OptiQ-4bit

Base model

Qwen/Qwen3.6-27B
Quantized
(12)
this model