Bonsai 27B oQ4e S32 Calibration Smoke

Experimental calibration artifact. Not a quality release. Do not use this model to judge Bonsai quality or as a production checkpoint.

Experiment family and current status

This is one public checkpoint in an unfinished MLX/oMLX compatibility and quantization experiment:

The experiment remains incomplete until maintained MLX/MLX-LM/oMLX support can load and generate through the relevant Bonsai paths without the current local compatibility patches, explicit calibration proxies, or imatrix-boundary workarounds. The larger predeclared evaluations follow that runtime gate.

The Hub's displayed parameter count reflects packed quantized storage tensors. It does not describe a newly trained smaller model; the logical source architecture is Bonsai 27B.

This repository preserves a bounded oQe-enhanced 4-bit MLX artifact that loaded and generated normally on a 32 GB Apple Silicon host. Its purpose is reproducible pipeline evidence and archival storage, not a claim that it outperforms another Bonsai checkpoint.

What this proves

  • An oQe-enhanced 4-bit MLX artifact was produced from the public, config-repaired BF16 conversion baseline.
  • Structural and load checks passed.
  • The 32-sample imatrix collection completed with 496 entries.
  • The allocation uses an effective 4.70 bits per weight: embeddings at 8-bit and 22 sensitivity-selected modules at 5-bit.
  • The artifact produced coherent reasoning in the bounded behavioral smoke.
  • On fixed 10-question screens it scored HellaSwag 6/10 and ARC-Challenge 6/10.

What this does not prove

The sensitivity stage was intentionally minimal: 2 samples of 64 tokens. The 10-question screens are too small for a quality ranking. This artifact has not passed the predeclared 100-question evaluations or a quality-sized sensitivity pass, and it did not beat the smaller oQ2e smoke artifact on the tiny screen.

Do not present it as a production model, a quality release, or evidence that oQ4e is superior to ordinary Q4, oQ2e, Q1, or another quantization method.

Reproduction parameters

Setting Value
oQ level 4
effective allocation 4.70 bpw
default quantization affine 4-bit, group size 64
embedding quantization affine 8-bit, group size 64
sensitivity-selected modules 22 at affine 5-bit
imatrix samples 32
imatrix sequence length 512
sensitivity samples 2
sensitivity sequence length 64
calibration dataset oqe_code_multilingual
imatrix entries 496

oq_imatrix_report.json, PROVENANCE.json, and the included IMATRIX_CACHE_S32.npz preserve the emitted artifact and calibration evidence. The NPZ contains aggregate activation statistics, not raw calibration prompts, and is not needed to run the model.

Bounded evaluation

The fixed screens used deterministic, thinking-disabled decoding and the same letter parser as the earlier oQ2e smoke:

Benchmark Score Peak generation memory
HellaSwag, fixed 10 6/10 15.70 GB
ARC-Challenge, fixed 10 6/10 15.42 GB

The raw records are retained under docs/. These are pipeline screens, not statistically useful benchmark claims.

Provenance

  • Derived baseline: TiGa-RCE/Bonsai-27B-MLX-BF16-Config-Repaired
  • Original source: prism-ml/Bonsai-27B-gguf
  • Source revision used for conversion: 0cf7e3d21581b169b4df1de8bf01316000e2fbb7
  • Original source file: Bonsai-27B-F16.gguf
  • Original source file SHA-256: d4a381a6d07131c34af888607bdbda49fc885c97673a0d22aa3e0f0284bba566
  • Output hashes: MODEL_SHA256SUMS.txt

The project retains the upstream Apache-2.0 LICENSE.txt and NOTICE.txt.

Runtime

The artifact uses heterogeneous MLX quantization metadata. It was validated with the experimental oMLX/oQe path used to create it. Compatibility with stock MLX-LM or other runtimes is not claimed.

Attribution

Created by Technologies Brewster Jennings du Canada for the Bonsai MLX quantization experiment. Created using Bonsai by Prism ML, derived from Qwen3.6-27B.

Downloads last month
60
Safetensors
Model size
4B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TiGa-RCE/Bonsai-27B-oQ4e-S32-Smoke

Base model

Qwen/Qwen3.6-27B
Quantized
(2)
this model