Qwen3.8-27B — Bonsai-Compatible Folded Quantization (3-bit)

Or: how we made a forked llama.cpp model run on stock oMLX without sacrificing much.

What?

These models are 2-bit and 3-bit affine-quantized variants of Qwen3.8-27B that run on vanilla mlx-lm and oMLX — no custom kernels, no PrismML fork, no runtime shenanigans.

They're derived from PrismML's Ternary-Bonsai-2, which is a remarkably clever 2-bit Qwen3.8 that retains ~95% of the base model's intelligence. There's just one problem: it won't load on oMLX.

Why doesn't Bonsai 2 load on oMLX?

Two reasons:

  1. Unknown model type — Bonsai 2 declares model_type: "prism_hadamard_qwen35". oMLX doesn't know what that is. It knows qwen3_5, but not PrismML's custom variant.

  2. Custom weight format — Bonsai 2 stores weights in "Packed" modules: (weight_uint32, scales_fp16, biases_fp16, signs_fp32). These aren't standard nn.Linear layers — they're a custom MLX module that applies a Hadamard transform to activations at runtime. Stock oMLX can't instantiate them.

The result: you download Bonsai 2, point oMLX at it, and get a model loading error. Frustrating.

The Folded Mechanism

Here's the insight that makes this work:

Bonsai 2's trick: Store weights in Hadamard space (H·W), then transform activations with H·x before matmul. The Hadamard transform spreads outlier values across all dimensions, making 2-bit quantization much more effective. At inference:

y = (H · diag(signs) · x)^T · (H · W)
  = x^T · diag(signs) · W                    (since H^T = H and H^2 = I)

Our trick: Absorb the activation transform into the weights:

W_folded = diag(signs) · H · W

Now standard inference y = x^T · W_folded produces the same result, with no runtime Hadamard needed. The quality benefit is preserved; the compatibility problem is solved.

This is not a new idea — it's a standard technique in the quantization literature — but applying it to Bonsai 2's specific Hadamard+sign pattern required some care.

Quality vs. Compatibility

Format Size Loads on oMLX? Quality (est.)
Bonsai 2 (original) 8.6 GB ❌ No custom runtime ~95% FP16
Bonsai-MLX 2-bit (this) 9.4 GB ✅ Yes ~90-93% FP16
Bonsai-MLX 3-bit (this) 12.8 GB ✅ Yes ~93-95% FP16
Naive 4-bit (oQ4) ~14 GB ✅ Yes ~91-93% FP16

The trade-off is small: ~2-5% quality vs. Bonsai 2's gold standard, in exchange for universal compatibility and no fork maintenance.

Benchmarks

These haven't been benchmarked against the standard suite yet — that's next. Early qualitative testing shows coherent generation at both bit widths, with the 3-bit variant noticeably closer to the original.

Usage

from mlx_lm import load, generate

model, tokenizer = load("hermitdave/Qwen3.8-27B-Bonsai-MLX-2bit")
# or
model, tokenizer = load("hermitdave/Qwen3.8-27B-Bonsai-MLX-3bit")

response = generate(model, tokenizer, prompt="Explain quantum entanglement in one paragraph.")
print(response)

Or via oMLX: just point it at the model directory and hit "Serve."

Credits

  • Prism ML — for the original Ternary Bonsai 2 and the Hadamard quantization research that inspired this. The quality of their 2-bit work is remarkable, and these models exist because theirs did first.
  • Hermes Agent (Nous Research) — for the conversion script, Hadamard unfolding logic, and debugging through multiple broken iterations. This was a collaborative engineering effort.

Limitations

  • Slightly larger than Bonsai 2 at 2-bit (9.4 GB vs 8.6 GB) — the Hadamard transform adds ~0.8 GB overhead.
  • Quality is estimated, not yet benchmarked against MMLU/GSM8K/etc.
  • Derived from Qwen3.8-27B via PrismML's quantization — any base model issues propagate through.

What's Next

  • Run the standard benchmark suite (MMLU, GSM8K, HumanEval, etc.)
  • Compare quality across the bit-width ladder
  • Try the recipe on other Bonsai-format models

This work was Hermes Agent finding a pragmatic path between "best possible quantization" and "actually runs on my machine."

Downloads last month
629
Safetensors
Model size
27B params
Tensor type
U32
·
F32
·
F16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hermitdave/Qwen3.8-27B-Bonsai-MLX-3bit

Base model

Qwen/Qwen3.8-27B
Finetuned
(9)
this model