Qwen3.8-27B-oQ2e-mtp

This model was quantized using oQ (oMLX v0.5.7) mixed-precision quantization.

Base model: Qwen/Qwen3.8-27B

Chat template: froggeric/Qwen-Fixed-Chat-Templates

Quantization details

  • Model type: qwen3_5
  • Bits: 2
  • Group size: 64
  • Format: MLX safetensors
  • MTP: Preserved (mtp_num_hidden_layers: 1)
  • Calibration: oQ2e (enhanced, imatrix-based; oqe_code_multilingual)

Environment

  • Hardware: M5 MacBook Air 32GB
  • Inference Framework: oMLX v0.5.7
  • Max Concurrent Requests: 4
  • Settings:
    • Thinking: Disabled
    • TurboQuant KV Cache: Enabled (4-bit)
    • Lightning MTP: Enabled (key speed improvement)

Performance Benchmarks

Note: Results are for reference only and may vary depending on hardware, software configuration, and workload.

Single Request Results

Test TTFT(ms) TPOT(ms) pp TPS tg TPS E2E(s) Throughput Peak Mem
pp1024/tg128 6074.9 67.08 168.6 tok/s 15.0 tok/s 14.615 78.8 tok/s 11.71 GB
pp4096/tg128 44281.6 96.54 92.5 tok/s 10.4 tok/s 56.590 74.6 tok/s 13.21 GB

Continuous Batching (pp1024 / tg128)

Batch tg TPS Speedup pp TPS pp TPS/req TTFT(ms) E2E(s)
1x 15.0 tok/s 1.00x 168.6 tok/s 168.6 tok/s 6074.9 14.615
2x 11.2 tok/s 0.75x 142.4 tok/s 71.2 tok/s 14382.5 37.157
4x 13.1 tok/s 0.87x 129.4 tok/s 32.4 tok/s 31210.7 70.791

Intelligence Benchmark

Note: Each benchmark round tests only 30 questions. Results are for reference only.

Benchmark Accuracy Correct Total Time(s) Think
MMLU 56.7% 17 30 189.5 No
TRUTHFULQA 63.3% 19 30 77.2 No
GSM8K 80.0% 24 30 649.4 No
MATHQA 23.3% 7 30 396.5 No
HUMANEVAL 83.3% 25 30 1468.4 No
Downloads last month
231
Safetensors
Model size
4B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including mlx-works/Qwen3.8-27B-oQ2e-mtp