Ornith-1.5-9B-oQ4e-mtp

This model was quantized using oQ (oMLX v0.6.2) mixed-precision quantization.

Base model: ornith-ai/Ornith-1.5-9B (Qwen3.5-9B based, RL self-improvement fine-tune; MTP head grafted from Qwen/Qwen3.5-9B)

Chat template: froggeric/Qwen-Fixed-Chat-Templates

Quantization details

  • Model type: qwen3_5
  • Bits: 4 (mixed-precision; linear-attention projections lifted to 5-bit)
  • Group size: 64
  • Mode: affine
  • Format: MLX safetensors
  • MTP: Preserved (mtp_num_hidden_layers: 1)
  • Calibration: oQ4e (enhanced, imatrix-based; oqe_code_multilingual)

Environment

  • Hardware: M5 MacBook Air 32GB
  • Inference Framework: oMLX v0.6.2
  • Max Concurrent Requests: 4
  • Settings:
    • Thinking: Disabled
    • TurboQuant KV Cache: Enabled (4-bit)
    • Lightning MTP: Enabled (key speed improvement)

Performance Benchmarks

Note: Results are for reference only and may vary depending on hardware, software configuration, and workload.

Both configurations: no thinking, TurboQuant KV 4-bit, Code (Python) context. Lightning MTP rows measured with MTP enabled; No MTP rows without.

Single Request Results

Test TTFT(ms) TPOT(ms) pp TPS tg TPS tg TPS (No MTP) E2E(s) Throughput Peak Mem
pp1024/tg128 1473.5 29.82 695.0 tok/s 33.8 tok/s 25.4 tok/s 5.308 217.1 tok/s 6.95 GB
pp4096/tg128 5923.1 31.49 691.5 tok/s 32.0 tok/s 24.7 tok/s 9.946 424.7 tok/s 7.69 GB

Continuous Batching (pp1024 / tg128)

Batch tg TPS tg TPS (No MTP) Speedup pp TPS pp TPS/req TTFT(ms) E2E(s)
1x 33.8 tok/s 25.4 tok/s 1.00x 695.0 tok/s 695.0 tok/s 1473.5 5.308
2x 38.6 tok/s 44.8 tok/s 1.14x 462.5 tok/s 231.3 tok/s 3332.9 11.057
4x 53.5 tok/s 61.1 tok/s 1.58x 385.1 tok/s 96.3 tok/s 6340.4 20.215

Intelligence Benchmark

Note: Each benchmark round tests only 30 questions. Results are for reference only.

Benchmark Accuracy Correct Total Time(s) Think
MMLU 66.7% 20 30 42.0 No
TRUTHFULQA 76.7% 23 30 13.4 No
GSM8K 86.7% 26 30 195.1 No
MATHQA 30.0% 9 30 27.6 No
HUMANEVAL 76.7% 23 30 192.1 No
Downloads last month
1,179
Safetensors
Model size
2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including mlx-works/Ornith-1.5-9B-oQ4e-mtp