M3 Ultra benchmark

#1
by ianphil - opened
MLX Community org

M3 Ultra benchmark (512 GB) β€” ~25 tok/s decode

Hardware: Apple M3 Ultra, 512 GB unified memory
Runtime: oMLX 0.6.3 (OpenAI-compatible server), model served as-is from this repo
Method: non-streaming chat completions via localhost; tok/s = usage.completion_tokens / wall time; max_tokens 400-512; default sampling

Test tok/s
Short prompt, reasoning_effort: low 24.6 / 25.6 (2 runs)
~1.5k-token prefill, incl. prefill 17.9
Short prompt, reasoning_effort: xhigh 24.2

Sign up or log in to comment