Bongard-mini

Machine intuition.

An open model trained on judgments, semantic relationships and action outcomes.

Bongard makes machine intuition a callable capability. Give it a situation, define the questions and possible outcomes, and receive a probability distribution for each question. The same model can read a customer conversation, compare an invoice with an order, or judge an agent's next action. It accepts text, structured data and images, and scores the supplied outcomes without generating text.

Read once. Judge in parallel.

Try the live demo Β· Code & runtime Β· Quick start Β· Design Β· Benchmarks Β· Examples

Intuition is learned

An experienced player sees a promising move. A reader catches irony. Such judgments draw on learned relationships, often before the person can explain each step. Their value extends beyond speed: a whole pattern can carry meaning that is hard to express as a list of reasons.

Bongard treats this form of judgment as an independent capability to design and train.

Stage What the model learns
Judgment Judge text, records and images. Learn which changes should alter an answer and which should leave it unchanged.
Relationships Connect judgments to their meaning across words, images and states through joint-embedding training.
Outcomes Learn what follows an action. Sandbox rollouts and exact solvers supply outcome distributions for candidate actions.

All three stages are complete. They update 7.09 billion trainable parameters, including the text backbone, multimodal projector and judgment head. The vision tower stays frozen. Each stage ran on one GPU: a B200 for stages 1 and 2, and a B300 for stage 3.

One reading, many decisions

Bongard-mini uses the T5 encoder–decoder structure of T5Gemma 2 4B-4B, with 7.51 billion total parameters. The encoder reads the evidence in both directions, with the question instructions in view. Separate decoder branches share that evidence. A trained head scores each candidate from its name and description. The runtime reads the state once and processes questions in groups of up to eight.

This structure serves a concrete purpose. A later clause can change the meaning of an earlier one; a log entry can explain an earlier failure. Bidirectional encoding lets those facts shape each other's representation before the model makes several judgments about the situation.

A wrapper changes the interface. Bongard trains the judgment.

OpenJev's default implementation wraps pretrained DiffusionGemma and reads answer-token probabilities. Kev trains LoRA adapters and a pointer head on Qwen, and reuses a causal state cache across questions. Bongard develops a shared, bidirectional evidence representation through full-model training on judgments, semantic relationships and action outcomes.

Jev is a trained, closed System One model. Bongard provides an open T5 route to the same class of judgments. The results below show how their strengths differ across tasks and measures. Full-model training costs more than adapter training; Bongard uses it to shape both the evidence representation and the judgment function.

Quick start

Use Python 3.12 or later. Install the runtime from GitHub and download the complete model bundle:

git clone --branch main --depth 1 https://github.com/AgentBull/bongard.git
cd bongard
python -m pip install .
hf download AgentBull/bongard-mini --local-dir bongard-mini

The examples below use CUDA. Use device="mps" on Apple Silicon or device="cpu" for CPU execution; the CLI accepts the same values with --device. The runtime loads the BF16 weights, with additional memory needed for the input and candidate descriptions.

Python: three questions, one state

import json
from bongard.inference import Predictor

predictor = Predictor.load(
    "bongard-mini",
    device="cuda",
    temperatures="bongard-mini/temperatures.json",
)

request = {
    "state": {
        "message": "My order arrived yesterday with a broken screen. Please replace it.",
        "policy": "Replace items damaged on delivery if reported within 14 days."
    },
    "questions": {
        "replace": {
            "type": "noul",
            "instructions": "Is this request eligible for a replacement under the policy?"
        },
        "route": {
            "type": "choice",
            "instructions": "Which team should handle this request?",
            "criteria": {
                "returns": "Damaged items, returns and replacements.",
                "billing": "Payment failures and invoice questions.",
                "sales": "New orders and product recommendations."
            }
        },
        "urgency": {
            "type": "score",
            "instructions": "How urgently does the customer need a response?",
            "criteria": [
                "Routine: can wait several days.",
                "Important: respond within one business day.",
                "Critical: immediate assistance is needed."
            ]
        }
    }
}

answers = predictor.predict(request)["answers"]
print(json.dumps(answers, indent=2))
print(answers["route"]["choice"])

Each entry in answers has the question ID you supplied:

Type How to define it How to read the answer
noul A yes/no question; optional descriptions of true and false noul is the probability of true, from 0 to 1.
choice A map of 2–255 candidate names to descriptions choice is the most likely candidate; probabilities contains the distribution over all candidates.
score A list of 2–10 ordered level descriptions, lowest first score is the expected zero-based level; probabilities uses keys "0", "1", etc.; legend maps them to descriptions.

Keep shared facts in state and put each decision in questions so the model can reuse its reading of the state. Use the supplied temperatures.json in both Python and the CLI to apply probability calibration. Loading only backbone/ with Transformers does not load the complete judgment model.

CLI and HTTP API

The GitHub repository includes a ready-to-run example request:

bongard predict --checkpoint bongard-mini --device cuda \
  --temperatures bongard-mini/temperatures.json \
  --request examples/request.json

To serve the model locally:

bongard serve --checkpoint bongard-mini --device cuda \
  --temperatures bongard-mini/temperatures.json \
  --host 127.0.0.1 --port 8000

In another terminal, from the same repository directory:

curl http://127.0.0.1:8000/v1/models

curl http://127.0.0.1:8000/v1/systemone \
  -H 'Content-Type: application/json' \
  --data-binary @examples/request.json

POST /v1/systemone uses the same request structure as Python and returns model, answers and usage. The API follows the System One wire format; usage.output_tokens is zero.

Using images

Python example

Add a top-level images array of base64 data URLs and refer to them as image 1, image 2, etc. PNG, JPEG and static WebP are supported. Local paths and remote URLs must be read and encoded first. For example, with your own package.png and the predictor loaded above:

import base64
from pathlib import Path

encoded = base64.b64encode(Path("package.png").read_bytes()).decode("ascii")
response = predictor.predict({
    "state": "A customer submitted a photo of their delivery.",
    "images": [f"data:image/png;base64,{encoded}"],
    "questions": {
        "damaged": {
            "type": "noul",
            "instructions": "Does image 1 show visible damage to the package?"
        }
    }
})
print(response["answers"]["damaged"]["noul"])

Benchmarks

Release evaluation: 2026-09-29, on one NVIDIA RTX PRO 6000 in BF16. These are our measurements on the splits and protocols below.

Benchmark Evaluation scope Bongard-mini
DecisionBench 1.0 All 23,900 decisions, 43 tasks, 2–255 candidates 78.3% accuracy; 78.0% on the 20,959 rows with no detected source-text overlap with training data
JevBench v1.4 231 public decisions: easy / standard / hard 100% / 97.2% / 57.7% accuracy
typed-decisions Test split: 400 cases, 2,000 decisions; zero-shot workflows 59.4% accuracy, 0.518 soft accuracy, 0.256 KL, 0.132 Brier
ImajevBench v2.0-lite Dev + calibration: 254 items with public gold 66.9% accuracy; chance 26.5%
behavior-benchmark Every trace; present / absent / not observable 0.458 / 0.668 F1 for present, core / multilingual
Fair-chance probes 18 kinds, including coins, dice, lotteries and tie-breaks 0.0009 mean KL from the uniform distribution; lower is better

How to read these results: DecisionBench's 78.3% result answers every row after increasing the request token budget for the 251 longest rows. At the default 64k-token budget, those rows fail and count as incorrect: accuracy is 77.4%. The overlap-filtered 78.0% result uses full coverage. JevBench scores cover its public questions. For typed-decisions, the gold is a teacher probability distribution, so its metrics measure agreement with that teacher. ImajevBench uses direct option scoring with an unknown option and averages all cyclic option orders.

DecisionBench comparison

DecisionBench 1.0 accuracy comparison across selected models

Comparison snapshot from the DecisionBench leaderboard, 2026-09-29. Reference model scores are published leaderboard results; Bongard-mini is our full-coverage evaluation on the same rows and gold labels, with the token-budget and overlap qualifications above.

Bongard and Jev, by task and measure

Evaluation Bongard-mini Jev 1.13
DecisionBench accuracy, all 23,900 decisions 78.3% 72.0%
typed-decisions top-label agreement, 2,000 decisions 59.4% 72.7%
typed-decisions KL from the teacher, lower is better 0.256 1.442
typed-decisions Brier score, lower is better 0.132 0.148

Jev matches the teacher's top label more often on typed-decisions. Bongard's full distributions are closer to the teacher's under KL and Brier. Reference scores come from the dataset cards and leaderboards on 2026-09-29. Bongard's per-item predictions let you inspect the decisions behind the scores.

Inference speed

Workload Measured result
Short JevBench easy/standard requests, one at a time (120 requests) 36.1 ms median, 39.3 ms p95 per decision
32 questions about one shared state 221 ms total, about 6.9 ms per decision
Independent JevBench requests, batch size 64 112.1 decisions/s, about 403,560 decisions/hour

These are local inference measurements and exclude network overhead. The batched result comes from batched model evaluation; the supplied HTTP server processes one request at a time. Input length, candidate descriptions and hardware affect latency.

To evaluate your server on typed-decisions, use the included benchmark runner from the repository root while the server is running:

python benchmarks/typed_decisions.py \
  --endpoint http://127.0.0.1:8000 --output typed-results.json

Judgments across situations

These are selected test scenarios from typed-decisions, with Bongard's published predictions. Each row shows one question from a five-question request.

Evidence Judgment Model output
A customer reports receiving the wrong item and asks what to do next. What is the conversation about? delivery: 71.0%
An order lists ten units; five were delivered and invoiced. Does the invoice reconcile with the order and delivery? false: 77.7%
An agent's certificate-rotation trace has eleven steps and no reported tool errors or constraint violations. What should the observability system do? continue: 66.9%

Game recordings

These selected gameplay recordings show Bongard-mini choosing actions from candidate lists. For board games, the engine supplies legal moves and their computed features; 2048 and Tetris also include engine ratings of moves. A short instruction states the strategy, and the model selects each action. The scores describe these selected recordings with the supplied engine features.

2048 β€” 4096 tile, 55,140 final score Tetris β€” consecutive line clears
Bongard-mini playing 2048 with a 4096 tile Bongard-mini choosing Tetris placements and clearing lines
Othello β€” 13–0 win against a positional opponent Snake β€” grows to 32 cells on an 8Γ—8 board
Bongard-mini eliminating the opponent's discs in Othello Bongard-mini playing Snake and reaching length 32

Doom β€” Defend the Center. The model receives a text description of the scene every 0.2 seconds and chooses turn_left, turn_right or fire. Across 32 episodes capped at 30 seconds, it averaged 10.25 kills, with a best of 20. The GIF shows an excerpt, with probabilities beside the action.

Bongard-mini choosing turns and shots in Doom from text scene descriptions

Evaluation notes

  • Worked-answer checking: in the evaluated tasks, the model rarely rejected incorrect worked solutions.
  • Calibration: use the supplied temperatures and evaluate probabilities on representative data. Refit calibration for a new deployment domain when needed.
  • Candidate order: Choice probabilities can depend on option order. rotations=True in Predictor.load, or --rotations in the CLI, averages cyclic option orders at extra compute cost; it is off by default.

Model and license

This is the final three-stage Bongard-mini checkpoint. The public GitHub repository contains the inference runtime, HTTP server, supervised training loop and a typed-decisions benchmark runner. The training data, joint-embedding objectives, sandbox environments and full training recipes are not part of the code release.

Model weights are subject to the Gemma Terms of Use; see NOTICE. The runtime code is Apache-2.0.

Cite

@techreport{ding2026bongard,
  title  = {Bongard: Training Machine Intuition},
  author = {Ding, Li and Jin, Haidi and Ji, Chen},
  institution = {AgentBull Pte Ltd},
  year   = {2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for AgentBull/bongard-mini

Finetuned
(7)
this model

Space using AgentBull/bongard-mini 1