Working great on M5 Max 128GB — thanks!

#1
by ericzhu0127 - opened

Thank you for this quant — it runs beautifully out of the box with mlx-serve.

My setup: M5 Max, 128 GB unified memory, serving via mlx-serve with a 262K context window.

Real-world results:

Decode: ~70 tok/s at 7K context, ~54 at 28K, ~42 at 118K (official instruct sampling: temp 0.7 / top_p 0.8 / top_k 20); 90+ tok/s with greedy

Prefill: ~1100 tok/s at short context, still ~450 tok/s at 118K

Long agentic workloads are solid: 90K+ token prompts with multiple tool calls, steady 65–70 tok/s decode

MTP speculative decoding works great: ~3.6–4 accepted drafts per round, 90–100% draft hit rate

Tool calling has been reliable across many consecutive agent loops, no repetition or drift

Appreciate you sharing this — it's become my daily driver for local agent work.
Codex 图像 2026年9月1日 15_31_42

Sign up or log in to comment