Instructions to use VertexAIco/Bonsai-27b-MLX-Optimization with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use VertexAIco/Bonsai-27b-MLX-Optimization with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("VertexAIco/Bonsai-27b-MLX-Optimization") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use VertexAIco/Bonsai-27b-MLX-Optimization with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "VertexAIco/Bonsai-27b-MLX-Optimization"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "VertexAIco/Bonsai-27b-MLX-Optimization" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use VertexAIco/Bonsai-27b-MLX-Optimization with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "VertexAIco/Bonsai-27b-MLX-Optimization"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "VertexAIco/Bonsai-27b-MLX-Optimization" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VertexAIco/Bonsai-27b-MLX-Optimization", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use VertexAIco/Bonsai-27b-MLX-Optimization with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "VertexAIco/Bonsai-27b-MLX-Optimization"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default VertexAIco/Bonsai-27b-MLX-Optimization
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use VertexAIco/Bonsai-27b-MLX-Optimization with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "VertexAIco/Bonsai-27b-MLX-Optimization"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "VertexAIco/Bonsai-27b-MLX-Optimization" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Bonsai 27b MLX Optimization
Benchmark at a glance
M4 Mac mini · 16 GB memory · short prompt · model already loaded.
| Version | Output tokens/sec | Speed vs original | First token | Task checks |
|---|---|---|---|---|
| Original 1-bit Bonsai, stock-style settings | 19.7 | 1.00× | 1.62 s | Not separately checked |
| Optimization | 19.8 | 1.00×; no meaningful gain | 1.60 s | 15/18 |
| Fast | 22.0 | 1.12× | 1.36 s | 11/18 |
| UltraFastUnStable, revised | 24.6 | 1.25× | 1.35 s | 11/18, earlier same-candidate check |
Read this before comparing: these are three-run medians for the same 81-token input and 128 output tokens, with thinking off and a tiny warm-up. Loading time is excluded. The revised UltraFastUnStable was measured later with a different memory ceiling, so its ratio is a comparison of recorded results—not a fresh same-session, controlled speedup. Longer inputs and tool schemas take longer. A higher rate does not mean better or correct answers; task counts are not percentages of intelligence retained.
The retired 48-block experiment reached 79.5 tokens/sec / 0.427 s TTFT, but produced nonsense or blank lines. It is unusable, is not the current UltraFastUnStable download, and is kept only in its historical revision tag. Measurement details and limitations.
Quick tour: full model, simpler local setup
- What it is: all 64 original blocks and unchanged 1-bit weights. Start here if answer quality matters most. The benefit is the launch/setup package—not a demonstrated big tokens/sec increase.
- What to download: use Files and versions or the full-repository CLI command below. You need the code, tokenizer, config, and Prism runtime too; launch optimizations are not inside one weight file.
- How to try it: follow Get started, then send a short prompt. Interactive mode keeps the model loaded. For tools, your app must validate and execute the model's requests; no web browser or image generator is built in.
The full model, packaged for running locally on an Apple Silicon Mac. This edition keeps all 64 layers and all of Prism ML's original 1-bit weights. We changed the supporting code and launch settings—not the model's weights.
Choose this edition if you care most about reliable answers, following instructions, summarizing search results, and making tool calls. It is still a 1-bit model and can make mistakes; unchanged weights do not mean perfect answers.
What you actually get
- The original 1-bit Bonsai checkpoint, with no additional layer removal or re-quantization. Its weight file is about 5.13 GB.
- A chat runner that can keep the model loaded, so every new request does not need to reload it from disk.
- Thinking-off defaults for quicker responses, memory limits, and an optional compact conversation cache in the text runner.
- A local server with an OpenAI-style API for apps that support that interface. Tools still need to be supplied and executed by your app: this model does not include a web browser or an image generator.
How fast is it?
On a Mac mini with an M4 chip and 16 GB memory, we measured:
| What we measured | Result |
|---|---|
| Output generation speed | 19.8 tokens per second |
| Time until the first output token | 1.60 seconds |
| Peak memory reported by MLX | 4.71 GB |
These are medians of three runs: the same 81-token input, 128 generated tokens, thinking off, and greedy decoding. The model was already loaded and warmed up. Loading it took roughly another three seconds and is not included in the 1.60-second result. Longer inputs take longer to read, and other Macs may have different speeds. MLX memory is not total system memory.
Does this make stock Bonsai faster? Not meaningfully in our short-input test. The same original weights with standard MLX generation settings measured 19.7 tokens/second and 1.62 seconds to the first token. The differences are too small to claim a real speedup. This package's practical benefits are convenient launching, keeping the model loaded, and controlling memory—not a proven big increase in tokens/second.
On an 18-task local check, this edition passed 15/18. It missed two requests to ask for image details and one requested sentence count. This is a small task check, not an overall intelligence score. We did not measure how much quality the upstream 1-bit quantization loses versus higher-precision Bonsai.
Important: download the whole package
Open Files and versions and download the entire repository, including the weights, configuration, tokenizer, installer, and Python/command files. Downloading only
model.safetensors, or choosing an app's “Download MLX” option, does not install all the supporting code or the required runtime. The launch optimizations are not baked into a single weight file.
This is not a universal drag-and-drop model for every MLX app. Its 1-bit format requires Prism ML's MLX fork. The installer builds that runtime locally; the compiled runtime is not included in the download.
Get started
You need an Apple Silicon Mac, uv, and Apple's Metal toolchain. The installer
uses one compilation job, but building the runtime still takes time and space.
Download all files with Hugging Face's CLI (sign in first for private access):
hf download VertexAIco/Bonsai-27b-MLX-Optimization --local-dir Bonsai-27b-MLX-Optimization
cd Bonsai-27b-MLX-Optimization
./install_prism_mlx.command
.venv-1bit/bin/python run_bonsai_1bit.py --interactive
For an app connection, run ./start_bonsai_1bit_server.command. The server
listens only on your Mac at http://127.0.0.1:8081/v1. A supported client must
connect to that API; simply importing the weights will not apply these settings.
Credit and license
Based on Prism ML's Bonsai 27B 1-bit model. The original Apache 2.0 license and notice are included. This edition changes supporting code only and preserves the published model tensors.
- Downloads last month
- 311
1-bit