Qwythos-9B-v2-oQ5-mtp

This model was quantized using oQ (oMLX v0.5.1) mixed-precision quantization.

Quantization details

  • Model type: qwen3_5
  • Bits: 5
  • Group size: 64
  • Format: MLX safetensors

Environment

  • Hardware: M5 MacBook Air 32GB
  • Inference Framework: oMLX v0.5.1
  • Settings:
    • Thinking: Disabled
    • Chat template parameter: enable_thinking=false (forced)
    • TurboQuant KV Cache: Disabled (if enable: reduces speed)
    • Native MTP: Enabled (key speed improvement)

Performance Benchmarks

Note: Results are for reference only and may vary depending on hardware, software configuration, and workload.

Single Request Results

Test TTFT(ms) TPOT(ms) pp TPS tg TPS E2E(s) Throughput Peak Mem
pp1024/tg128 1654.0 27.87 619.1 tok/s 36.2 tok/s 5.224 220.5 tok/s 7.16 GB
pp4096/tg128 6545.4 27.50 625.8 tok/s 36.6 tok/s 10.072 419.4 tok/s 7.75 GB

Continuous Batching (pp1024 / tg128)

Batch tg TPS Speedup pp TPS pp TPS/req TTFT(ms) E2E(s)
1x 36.2 tok/s 1.00x 619.1 tok/s 619.1 tok/s 1654.0 5.224
2x 39.7 tok/s 1.10x 574.8 tok/s 287.4 tok/s 3562.8 10.005
4x 62.6 tok/s 1.73x 557.2 tok/s 139.3 tok/s 7169.9 15.529

Intelligence Benchmark

Note: Each benchmark round tests only 30 questions. Results are for reference only.

Benchmark Accuracy Correct Total Time(s) Think
MMLU 70.0% 21 30 38.5 No
TRUTHFULQA 93.3% 28 30 15.0 No
GSM8K 86.7% 26 30 136.0 No
MATHQA 43.3% 13 30 22.1 No
HUMANEVAL 80.0% 24 30 109.0 No

Summary

Benchmark Accuracy
MMLU 70.0%
TRUTHFULQA 93.3%
GSM8K 86.7%
MATHQA 43.3%
HUMANEVAL 80.0%
Average 74.7%
Downloads last month
159
Safetensors
Model size
2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

5-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including mlx-works/Qwythos-9B-v2-oQ5-mtp