Heuristic 1
One model out of two: Bonsai 2 27B (ternary) thinks, Decider-4B judges, MetaCog wires them into a single OpenAI-compatible endpoint.
This is not a weight merge. No tensors were averaged. It is an inference-time pipeline merge: the thinker writes candidate answers, an independent System-One judge (Decider-4B) scores them, and the winner is returned. One endpoint, one model ID, two models inside.
user prompt βββΆ bonsai-27B βββΆ greedy anchor judged FIRST
β β₯ 0.80 and finished? βββΆ answer immediately
βΌ
7 sampled paths IN PARALLEL,
each judged the moment it finishes;
first finished path β₯ 0.80 wins on the spot,
the rest are abandoned
β
(nothing clears the bar? best judged
finished path wins)
The judge is not the thinker β that separation is the whole point.
race.py implements this loop (parallel generation, judge-as-they-finish,
early stop). The old sequential MetaCog best_of_n is still available via
NO_RACE=1 (server) or --no-race (batch) for A/B comparison. The
early-stop bar is STOP_CONFIDENCE / --stop-confidence (default 0.80).
On the 0.80 bar, honestly: the MetaCog repo measured its 0.95 cascade threshold on real Jev (a noul β₯ 0.95 was correct 226/229 times). 0.80 did not win anywhere β it is the operator's directive and is untested on decider-4b, whose confidence scale is not calibrated (values above 1.0 have been observed). Measure before trusting it.
Measured
| eval | bonsai alone | merge (Heuristic 1) | note |
|---|---|---|---|
HumanEval, n=164, true pass@1 (prompt+completion graded as one program, <think> stripped) |
108/164 (65.9%) | 135/164 (82.3%) | judge fixed 27, broke 0; McNemar p < 1e-6 |
| 12 easy Q&A (verified) | 12/12 | 12/12 | tie; nothing to fix |
| 10 hard Q&A (verified, tricks + big arithmetic) | 8/10 | 10/10 | judge fixed 2, broke 0 |
Cost per problem (HumanEval averages): merge uses ~5,600 thinker tokens over 2.0 thinker calls + 8.5 judge calls; bonsai alone uses ~700 tokens in 1 call. Roughly 8Γ the thinker tokens for +16 points of accuracy, $0 on local hardware.
Latency: when the judge is confident in the greedy answer it returns
immediately (~same as bonsai alone β e.g. 2.5s on a trivial prompt); the
full vote only runs when uncertain. Measured on 3 HumanEval problems where
the full vote ran (DGX Spark, sequential paths, old loop): heuristic-1 took
210s / 406s / 440s vs bonsai alone at 17s / 18s / 18s β roughly 12β24Γ slower
wall-clock on hard problems. The race loop (race.py) parallelizes the
sampled paths and cuts off the moment any finished path clears 0.80, which
removes most of that gap; re-measure on the Spark to confirm. Every response
carries merge.judge_confidence.
Run it
pip install -r requirements.txt # metacog + httpx + pydantic
# one model on :8200
python server.py
# POST /v1/chat/completions {"messages": [...], "model": "bonsai-decider4b-merge"}
# batch runner
python merge.py --problem "What is 17*23?"
python merge.py --problems problems.jsonl --out results.jsonl
Env knobs: THINKER_URL (default http://localhost:8010),
THINKER_MODEL (default bonsai), JUDGE_URL (default
http://localhost:8008), PORT (8200), N_PATHS (8), NO_CASCADE,
MAX_TOKENS (2048), TEMPERATURE (0.8).
You need both backends up:
- Thinker β any OpenAI-compatible chat endpoint serving Bonsai 2 27B.
Reference setup: bjev's PrismML llama.cpp fork on
:8010(Ternary-Bonsai-2-27B-PQ2_0.gguf, 6.66 GiB). - Judge β a
/v1/systemoneendpoint serving Decider-4B. Reference:runs/judge_server.py(a/v1/systemoneshim over theMapika/decider-4bcheckpoint) on:8008.
On a DGX Spark this fits alongside both models in ~20 GB and runs at $0.
Files
server.pyβ the one-model OpenAI-compatible server (:8200)merge.pyβ batch runner (problems in, answers + confidence + timings out)eval/ab_bonsai_vs_merge.pyβ A/B harness: bonsai alone vs the mergeeval/problems_easy.jsonl,eval/problems_hard.jsonlβ the measured sets
License
MIT. This repo ships pipeline code only β no weights. Parent model licenses (checked 2026-09-27):
- Bonsai 2 27B ternary (
prism-ml/Ternary-Bonsai-2-27B-gguf) β Apache-2.0 - Decider-4B (
Mapika/decider-4b) β Apache-2.0 - MetaCog β MIT; llama.cpp β MIT Fetch the weights yourself under their Apache-2.0 terms; nothing here relicenses them. Not affiliated with PrismML, Mapika, or Alibaba.