YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

MLX Quantisation Ladders β€” Measurements

Five uncensored model families, each converted from one BF16 source with one group size, so bit width is the only variable within a family. Converted with mlx_vlm.convert, measured on a single Apple M3 Ultra (96 GB unified, macOS 27).

Perplexity: allenai/tulu-3-sft-mixture, 192 samples x 512 tokens, seed 123 β€” identical for every rung. Throughput: aggregate output tok/s on ~70-token prompts, with a unique nonce per request so no measurement reads another's prefix cache (cached_tokens asserted 0 on every level).

Perplexity is comparable only WITHIN a family. Tokenizers differ across families (Gemma 262,144 vocab vs Qwen 151,936), so a cross-family comparison of these numbers is meaningless. The x best column is the portable one.

Muse-Glimmer-30B-Abliterated

rung size GB perplexity x best status
q3 15.95 8.776 1.21x published
q4 19.44 7.300 1.01x published
q5 22.93 7.306 1.01x published
q6 26.42 7.284 1.01x published
q8 33.4 7.242 1.00x published
mxfp4 18.57 7.545 1.04x published
mixed36 18.27 8.336 1.15x published

Qwen3.8-27B-Abliterated

rung size GB perplexity x best status
q3 12.72 7.178 1.15x published
q4 16.08 6.455 1.03x published
q5 19.44 6.419 1.03x published
q6 22.8 6.427 1.03x published
q8 29.53 6.452 1.03x published
mxfp4 15.24 6.251 1.00x published
mixed36 14.76 6.807 1.09x published

Ornith-1.5-9B-Abliterated

rung size GB perplexity x best status
q3 5.34 7.518 1.41x published
q4 6.46 5.471 1.03x published
q5 7.58 5.325 1.00x published
q6 8.7 5.355 1.01x published
q8 10.94 5.333 1.00x published
mxfp4 6.18 5.783 1.09x published
mixed36 6.42 6.736 1.26x published

Gemma-4-12B-Abliterated

rung size GB perplexity x best status
q3 5.28 136393.505 976.16x withheld
q4 6.77 287.296 2.06x withheld
q5 8.27 211.862 1.52x published
q6 9.76 144.271 1.03x published
q8 12.75 139.724 1.00x published
mxfp4 6.4 254.121 1.82x published
mixed36 6.24 31094.364 222.54x withheld

Gemma-4-26B-A4B-Heretic

rung size GB perplexity x best status
q3 12.22 547.455 5.45x withheld
q4 15.37 168.023 1.67x published
q6 21.68 100.433 1.00x published
q8 27.99 105.975 1.06x published
mxfp4 14.59 195.105 1.94x published
mixed36 13.87 290.832 2.90x withheld

What the numbers say

4-bit is free on the Qwen lineage. Qwen3.8-27B, Ornith-1.5-9B and Muse-Glimmer-30B all sit within ~1.03x of their best rung from 4-bit upward. 8-bit costs 70%+ more disk for no measurable quality gain.

Gemma is quantisation-fragile. Both Gemma-4 families degrade noticeably at 4-bit (1.67x and 2.06x) and collapse at 3-bit (5.5x and 976x) under the identical pipeline that leaves Qwen models flat. That is a property of the architecture, not of the tooling.

MXFP4 is often the best rung available. Best-in-family on Qwen3.8-27B (1.00x), and on Muse-Glimmer it was simultaneously the fastest rung measured and the second-smallest.

2-bit was built, measured and deleted. Across four families it ran 15x to 5e9x its family best while saving barely a gigabyte over 3-bit. No 2-bit artifacts are published.

Selection rule

A rung is published only if:

  1. its weight index exists and every shard it names is present and non-trivial, and
  2. its perplexity is <= 2x the best rung in its own family.

Five rungs were withheld. Each withheld rung's ratio appears in the tables above, so the decision is auditable rather than asserted.

Loading

These are mlx-vlm models. mlx_lm cannot load muse_glimmer or gemma4_unified at all β€” those architectures are registered only in mlx-vlm.

pip install mlx-vlm
mlx_vlm.generate --model shoemoney/<repo> --prompt "Hello" --max-tokens 256
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support