YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
MLX Quantisation Ladders β Measurements
Five uncensored model families, each converted from one BF16 source with one
group size, so bit width is the only variable within a family. Converted with
mlx_vlm.convert, measured on a single Apple M3 Ultra (96 GB unified, macOS 27).
Perplexity: allenai/tulu-3-sft-mixture, 192 samples x 512 tokens, seed 123 β
identical for every rung. Throughput: aggregate output tok/s on ~70-token
prompts, with a unique nonce per request so no measurement reads another's
prefix cache (cached_tokens asserted 0 on every level).
Perplexity is comparable only WITHIN a family. Tokenizers differ across families (Gemma 262,144 vocab vs Qwen 151,936), so a cross-family comparison of these numbers is meaningless. The
x bestcolumn is the portable one.
Muse-Glimmer-30B-Abliterated
| rung | size GB | perplexity | x best | status |
|---|---|---|---|---|
| q3 | 15.95 | 8.776 | 1.21x | published |
| q4 | 19.44 | 7.300 | 1.01x | published |
| q5 | 22.93 | 7.306 | 1.01x | published |
| q6 | 26.42 | 7.284 | 1.01x | published |
| q8 | 33.4 | 7.242 | 1.00x | published |
| mxfp4 | 18.57 | 7.545 | 1.04x | published |
| mixed36 | 18.27 | 8.336 | 1.15x | published |
Qwen3.8-27B-Abliterated
| rung | size GB | perplexity | x best | status |
|---|---|---|---|---|
| q3 | 12.72 | 7.178 | 1.15x | published |
| q4 | 16.08 | 6.455 | 1.03x | published |
| q5 | 19.44 | 6.419 | 1.03x | published |
| q6 | 22.8 | 6.427 | 1.03x | published |
| q8 | 29.53 | 6.452 | 1.03x | published |
| mxfp4 | 15.24 | 6.251 | 1.00x | published |
| mixed36 | 14.76 | 6.807 | 1.09x | published |
Ornith-1.5-9B-Abliterated
| rung | size GB | perplexity | x best | status |
|---|---|---|---|---|
| q3 | 5.34 | 7.518 | 1.41x | published |
| q4 | 6.46 | 5.471 | 1.03x | published |
| q5 | 7.58 | 5.325 | 1.00x | published |
| q6 | 8.7 | 5.355 | 1.01x | published |
| q8 | 10.94 | 5.333 | 1.00x | published |
| mxfp4 | 6.18 | 5.783 | 1.09x | published |
| mixed36 | 6.42 | 6.736 | 1.26x | published |
Gemma-4-12B-Abliterated
| rung | size GB | perplexity | x best | status |
|---|---|---|---|---|
| q3 | 5.28 | 136393.505 | 976.16x | withheld |
| q4 | 6.77 | 287.296 | 2.06x | withheld |
| q5 | 8.27 | 211.862 | 1.52x | published |
| q6 | 9.76 | 144.271 | 1.03x | published |
| q8 | 12.75 | 139.724 | 1.00x | published |
| mxfp4 | 6.4 | 254.121 | 1.82x | published |
| mixed36 | 6.24 | 31094.364 | 222.54x | withheld |
Gemma-4-26B-A4B-Heretic
| rung | size GB | perplexity | x best | status |
|---|---|---|---|---|
| q3 | 12.22 | 547.455 | 5.45x | withheld |
| q4 | 15.37 | 168.023 | 1.67x | published |
| q6 | 21.68 | 100.433 | 1.00x | published |
| q8 | 27.99 | 105.975 | 1.06x | published |
| mxfp4 | 14.59 | 195.105 | 1.94x | published |
| mixed36 | 13.87 | 290.832 | 2.90x | withheld |
What the numbers say
4-bit is free on the Qwen lineage. Qwen3.8-27B, Ornith-1.5-9B and Muse-Glimmer-30B all sit within ~1.03x of their best rung from 4-bit upward. 8-bit costs 70%+ more disk for no measurable quality gain.
Gemma is quantisation-fragile. Both Gemma-4 families degrade noticeably at 4-bit (1.67x and 2.06x) and collapse at 3-bit (5.5x and 976x) under the identical pipeline that leaves Qwen models flat. That is a property of the architecture, not of the tooling.
MXFP4 is often the best rung available. Best-in-family on Qwen3.8-27B (1.00x), and on Muse-Glimmer it was simultaneously the fastest rung measured and the second-smallest.
2-bit was built, measured and deleted. Across four families it ran 15x to 5e9x its family best while saving barely a gigabyte over 3-bit. No 2-bit artifacts are published.
Selection rule
A rung is published only if:
- its weight index exists and every shard it names is present and non-trivial, and
- its perplexity is <= 2x the best rung in its own family.
Five rungs were withheld. Each withheld rung's ratio appears in the tables above, so the decision is auditable rather than asserted.
Loading
These are mlx-vlm models. mlx_lm cannot load muse_glimmer or
gemma4_unified at all β those architectures are registered only in mlx-vlm.
pip install mlx-vlm
mlx_vlm.generate --model shoemoney/<repo> --prompt "Hello" --max-tokens 256