Heuristic 1

One model out of two: Bonsai 2 27B (ternary) thinks, Decider-4B judges, MetaCog wires them into a single OpenAI-compatible endpoint.

This is not a weight merge. No tensors were averaged. It is an inference-time pipeline merge: the thinker writes candidate answers, an independent System-One judge (Decider-4B) scores them, and the winner is returned. One endpoint, one model ID, two models inside.

user prompt ──▢ bonsai-27B ──▢ greedy anchor judged FIRST
                                  β”‚  β‰₯ 0.80 and finished? ──▢ answer immediately
                                  β–Ό
                        7 sampled paths IN PARALLEL,
                        each judged the moment it finishes;
                        first finished path β‰₯ 0.80 wins on the spot,
                        the rest are abandoned
                                  β”‚
                        (nothing clears the bar? best judged
                         finished path wins)

The judge is not the thinker β€” that separation is the whole point.

race.py implements this loop (parallel generation, judge-as-they-finish, early stop). The old sequential MetaCog best_of_n is still available via NO_RACE=1 (server) or --no-race (batch) for A/B comparison. The early-stop bar is STOP_CONFIDENCE / --stop-confidence (default 0.80).

On the 0.80 bar, honestly: the MetaCog repo measured its 0.95 cascade threshold on real Jev (a noul β‰₯ 0.95 was correct 226/229 times). 0.80 did not win anywhere β€” it is the operator's directive and is untested on decider-4b, whose confidence scale is not calibrated (values above 1.0 have been observed). Measure before trusting it.

Measured

eval bonsai alone merge (Heuristic 1) note
HumanEval, n=164, true pass@1 (prompt+completion graded as one program, <think> stripped) 108/164 (65.9%) 135/164 (82.3%) judge fixed 27, broke 0; McNemar p < 1e-6
12 easy Q&A (verified) 12/12 12/12 tie; nothing to fix
10 hard Q&A (verified, tricks + big arithmetic) 8/10 10/10 judge fixed 2, broke 0

Cost per problem (HumanEval averages): merge uses ~5,600 thinker tokens over 2.0 thinker calls + 8.5 judge calls; bonsai alone uses ~700 tokens in 1 call. Roughly 8Γ— the thinker tokens for +16 points of accuracy, $0 on local hardware.

Latency: when the judge is confident in the greedy answer it returns immediately (~same as bonsai alone β€” e.g. 2.5s on a trivial prompt); the full vote only runs when uncertain. Measured on 3 HumanEval problems where the full vote ran (DGX Spark, sequential paths, old loop): heuristic-1 took 210s / 406s / 440s vs bonsai alone at 17s / 18s / 18s β€” roughly 12–24Γ— slower wall-clock on hard problems. The race loop (race.py) parallelizes the sampled paths and cuts off the moment any finished path clears 0.80, which removes most of that gap; re-measure on the Spark to confirm. Every response carries merge.judge_confidence.

Run it

pip install -r requirements.txt   # metacog + httpx + pydantic

# one model on :8200
python server.py
# POST /v1/chat/completions {"messages": [...], "model": "bonsai-decider4b-merge"}

# batch runner
python merge.py --problem "What is 17*23?"
python merge.py --problems problems.jsonl --out results.jsonl

Env knobs: THINKER_URL (default http://localhost:8010), THINKER_MODEL (default bonsai), JUDGE_URL (default http://localhost:8008), PORT (8200), N_PATHS (8), NO_CASCADE, MAX_TOKENS (2048), TEMPERATURE (0.8).

You need both backends up:

  • Thinker β€” any OpenAI-compatible chat endpoint serving Bonsai 2 27B. Reference setup: bjev's PrismML llama.cpp fork on :8010 (Ternary-Bonsai-2-27B-PQ2_0.gguf, 6.66 GiB).
  • Judge β€” a /v1/systemone endpoint serving Decider-4B. Reference: runs/judge_server.py (a /v1/systemone shim over the Mapika/decider-4b checkpoint) on :8008.

On a DGX Spark this fits alongside both models in ~20 GB and runs at $0.

Files

  • server.py β€” the one-model OpenAI-compatible server (:8200)
  • merge.py β€” batch runner (problems in, answers + confidence + timings out)
  • eval/ab_bonsai_vs_merge.py β€” A/B harness: bonsai alone vs the merge
  • eval/problems_easy.jsonl, eval/problems_hard.jsonl β€” the measured sets

License

MIT. This repo ships pipeline code only β€” no weights. Parent model licenses (checked 2026-09-27):

  • Bonsai 2 27B ternary (prism-ml/Ternary-Bonsai-2-27B-gguf) β€” Apache-2.0
  • Decider-4B (Mapika/decider-4b) β€” Apache-2.0
  • MetaCog β€” MIT; llama.cpp β€” MIT Fetch the weights yourself under their Apache-2.0 terms; nothing here relicenses them. Not affiliated with PrismML, Mapika, or Alibaba.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for CommanderCuth/Heuristic-1

Finetuned
(3)
this model