Bonsai 27b MLX Optimization

Benchmark at a glance

M4 Mac mini · 16 GB memory · short prompt · model already loaded.

Version Output tokens/sec Speed vs original First token Task checks
Original 1-bit Bonsai, stock-style settings 19.7 1.00× 1.62 s Not separately checked
Optimization 19.8 1.00×; no meaningful gain 1.60 s 15/18
Fast 22.0 1.12× 1.36 s 11/18
UltraFastUnStable, revised 24.6 1.25× 1.35 s 11/18, earlier same-candidate check

Read this before comparing: these are three-run medians for the same 81-token input and 128 output tokens, with thinking off and a tiny warm-up. Loading time is excluded. The revised UltraFastUnStable was measured later with a different memory ceiling, so its ratio is a comparison of recorded results—not a fresh same-session, controlled speedup. Longer inputs and tool schemas take longer. A higher rate does not mean better or correct answers; task counts are not percentages of intelligence retained.

The retired 48-block experiment reached 79.5 tokens/sec / 0.427 s TTFT, but produced nonsense or blank lines. It is unusable, is not the current UltraFastUnStable download, and is kept only in its historical revision tag. Measurement details and limitations.

Quick tour: full model, simpler local setup

  1. What it is: all 64 original blocks and unchanged 1-bit weights. Start here if answer quality matters most. The benefit is the launch/setup package—not a demonstrated big tokens/sec increase.
  2. What to download: use Files and versions or the full-repository CLI command below. You need the code, tokenizer, config, and Prism runtime too; launch optimizations are not inside one weight file.
  3. How to try it: follow Get started, then send a short prompt. Interactive mode keeps the model loaded. For tools, your app must validate and execute the model's requests; no web browser or image generator is built in.

The full model, packaged for running locally on an Apple Silicon Mac. This edition keeps all 64 layers and all of Prism ML's original 1-bit weights. We changed the supporting code and launch settings—not the model's weights.

Choose this edition if you care most about reliable answers, following instructions, summarizing search results, and making tool calls. It is still a 1-bit model and can make mistakes; unchanged weights do not mean perfect answers.

What you actually get

  • The original 1-bit Bonsai checkpoint, with no additional layer removal or re-quantization. Its weight file is about 5.13 GB.
  • A chat runner that can keep the model loaded, so every new request does not need to reload it from disk.
  • Thinking-off defaults for quicker responses, memory limits, and an optional compact conversation cache in the text runner.
  • A local server with an OpenAI-style API for apps that support that interface. Tools still need to be supplied and executed by your app: this model does not include a web browser or an image generator.

How fast is it?

On a Mac mini with an M4 chip and 16 GB memory, we measured:

What we measured Result
Output generation speed 19.8 tokens per second
Time until the first output token 1.60 seconds
Peak memory reported by MLX 4.71 GB

These are medians of three runs: the same 81-token input, 128 generated tokens, thinking off, and greedy decoding. The model was already loaded and warmed up. Loading it took roughly another three seconds and is not included in the 1.60-second result. Longer inputs take longer to read, and other Macs may have different speeds. MLX memory is not total system memory.

Does this make stock Bonsai faster? Not meaningfully in our short-input test. The same original weights with standard MLX generation settings measured 19.7 tokens/second and 1.62 seconds to the first token. The differences are too small to claim a real speedup. This package's practical benefits are convenient launching, keeping the model loaded, and controlling memory—not a proven big increase in tokens/second.

On an 18-task local check, this edition passed 15/18. It missed two requests to ask for image details and one requested sentence count. This is a small task check, not an overall intelligence score. We did not measure how much quality the upstream 1-bit quantization loses versus higher-precision Bonsai.

Important: download the whole package

Open Files and versions and download the entire repository, including the weights, configuration, tokenizer, installer, and Python/command files. Downloading only model.safetensors, or choosing an app's “Download MLX” option, does not install all the supporting code or the required runtime. The launch optimizations are not baked into a single weight file.

This is not a universal drag-and-drop model for every MLX app. Its 1-bit format requires Prism ML's MLX fork. The installer builds that runtime locally; the compiled runtime is not included in the download.

Get started

You need an Apple Silicon Mac, uv, and Apple's Metal toolchain. The installer uses one compilation job, but building the runtime still takes time and space.

Download all files with Hugging Face's CLI (sign in first for private access):

hf download VertexAIco/Bonsai-27b-MLX-Optimization --local-dir Bonsai-27b-MLX-Optimization
cd Bonsai-27b-MLX-Optimization
./install_prism_mlx.command
.venv-1bit/bin/python run_bonsai_1bit.py --interactive

For an app connection, run ./start_bonsai_1bit_server.command. The server listens only on your Mac at http://127.0.0.1:8081/v1. A supported client must connect to that API; simply importing the weights will not apply these settings.

Credit and license

Based on Prism ML's Bonsai 27B 1-bit model. The original Apache 2.0 license and notice are included. This edition changes supporting code only and preserves the published model tensors.

Downloads last month
311
Safetensors
Model size
2B params
Tensor type
U32
·
F16
·
MLX
Hardware compatibility
Log In to add your hardware

1-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VertexAIco/Bonsai-27b-MLX-Optimization

Base model

Qwen/Qwen3.6-27B
Finetuned
(3)
this model

Collection including VertexAIco/Bonsai-27b-MLX-Optimization