Instructions to use emrevrg/AUBIN-31B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use emrevrg/AUBIN-31B with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("google/gemma-4-31B-it") model = PeftModel.from_pretrained(base_model, "emrevrg/AUBIN-31B") - Notebooks
- Google Colab
- Kaggle
- AUBIN-31B β AUBIN by Norovox
- At a glance
- Head-to-head with Jev β on Jev's own published split
- Where AUBIN wins by multiples (same questions)
- Against the Kev family (Kev README, development / test, clean)
- Computer use: Mind2Web web-agent benchmark
- Game: AUBIN plays 2048 (live replay)
- Real-time control: command-following grid game
- Consistency and hallucination ("gray area" test types)
- AUBIN Loop: generator + decider (
aubin loop) - Measured speed (free Kaggle GPUs, 4-bit, batch = 1 case)
- Use
- Honest notes
- AUBIN-Learn β instant self-learning (experimental, Norovox core)
- Results, 3 October 2026 (full report:
reports/AUBIN_RESULTS_2026-10-03.md)
- At a glance
AUBIN-31B β AUBIN by Norovox
At a glance
- Beats Jev on Jev's own published split, all 8 numbers: accuracy 0.863 vs 0.857 and 0.862 vs 0.845; NLL 0.409 vs 0.701; Brier, ECE lower too.
- AUBIN Trio (adds Phi-4), 8/8 vs Jev: accuracy 0.872 vs 0.857 (new sources) and 0.851 vs 0.845 (trained sources); ECE 0.018 / 0.032 vs 0.049 / 0.067.
- AUBIN Quad (adds Phi-4 + Mistral-Small; chosen by the rule), 8/8 vs Jev: accuracy 0.869 vs 0.857 (new sources) and 0.856 vs 0.845 (trained sources); ECE 0.030 / 0.022 vs 0.049 / 0.067.
- 22x fewer confidently-wrong answers (AUBIN-31B, one pass, trained-source split: 0.24% vs Jev 5.22%).
- 0 errors on held-out policy rules it never trained on (Jev: 14 / 176).
- Locked test, never-trained sources: 0.898, the best in Kev's table (Kev-27B 0.896).
- Computer use (Mind2Web), AUBIN-31B with no Mind2Web training beats GPT-4 on 9 of 9 numbers (element acc / op F1 / step success): Cross-Task 48.0 / 81.0 / 40.0 vs 41.6 / 60.6 / 36.2; Cross-Website 41.0 / 82.0 / 33.5 vs 35.8 / 51.1 / 30.1; Cross-Domain 54.0 / 86.5 / 48.5 vs 37.1 / 46.5 / 26.4. GPT-4 in the paper chose from 10 candidates, AUBIN from 50.
- Also beats the fine-tuned MindAct (Flan-T5-XL, trained on Mind2Web) on all three numbers: Cross-Domain (54.0 / 86.5 / 48.5 vs 42.1 / 66.5 / 39.6).
- Real-time control (command-following game, 100 unseen episodes, success rate): AUBIN-E4B-Control 95% with the safety shield (0 lava deaths, 385 ms/move on a free T4), 58% with no shield at all; AUBIN-12B-Control 92% with the safety shield (0 lava deaths, 856 ms/move on a free T4), 74% with no shield at all. Rule baselines: greedy 58%, greedy + same shield 89%.
- AUBIN Loop (generator + decider, same Gemma weights): GSM8K (400 problems) 84.2% β 92.5%; majority vote of the same samples 85.8%.
- Plays 2048 with a lookahead tool: mean 29,009, 2048 tile in 8/10 games (replay).
Typed questions in, calibrated probabilities out. AUBIN is Norovox's open-base model family: strong open-weight bases, adapted and calibrated for structured decisions (choice / yes-no / score). Fast path = one forward pass; when the model is unsure (top-2 log-prob gap < 4), it writes a short rationale and re-scores (think-when-uncertain, ~14% of questions). Open weights, runs locally/offline in 4-bit on 2x free 16 GB GPUs (Kaggle 2xT4).
Head-to-head with Jev β on Jev's own published split
Jev's published numbers (jaredpalmer/kev runs/jev-*-v4/report.json, "clean")
are on the development split of Kev's public suites, variant == clean (656 new-source + 1264 trained-source
questions). AUBIN is scored on exactly the same questions (code/jev_compare.py). Bold = better than Jev.
| model | new sources acc | Brier β | ECE β | NLL β | trained sources acc | Brier β | ECE β | NLL β |
|---|---|---|---|---|---|---|---|---|
| Jev (hosted, published report) | 0.857 | 0.2110 | 0.049 | 0.701 | 0.845 | 0.2367 | 0.067 | 0.776 |
| AUBIN-12B full precision (think) | 0.857 | 0.2331 | 0.050 | 0.460 | 0.851 | 0.2457 | 0.038 | 0.623 |
| AUBIN-31B (fast, one pass) | 0.852 | 0.2280 | 0.068 | 0.444 | 0.850 | 0.2365 | 0.063 | 0.567 |
| AUBIN-31B (think) | 0.863 | 0.2255 | 0.033 | 0.433 | 0.857 | 0.2300 | 0.040 | 0.569 |
| AUBIN Duo (31B + 12B, think) | 0.867 | 0.2137 | 0.021 | 0.417 | 0.859 | 0.2193 | 0.031 | 0.554 |
| AUBIN Duo (31B + 12B full precision, think) | 0.863 | 0.2114 | 0.035 | 0.411 | 0.858 | 0.2243 | 0.029 | 0.554 |
| AUBIN Duo-SC (31B self-consistent think + 12B full precision) | 0.863 | 0.2104 | 0.039 | 0.409 | 0.862 | 0.2227 | 0.028 | 0.551 |
| AUBIN Trio (31B-SC + 12B full precision + Phi-4) | 0.872 | 0.2087 | 0.018 | 0.406 | 0.851 | 0.2315 | 0.032 | 0.536 |
| AUBIN Quad (31B-SC + 12B full precision + Phi-4 + Mistral-Small) | 0.869 | 0.2068 | 0.030 | 0.404 | 0.856 | 0.2268 | 0.022 | 0.536 |
AUBIN Duo-SC (31B self-consistent think + 12B full precision) is better than Jev on all 8 numbers above (accuracy, Brier, ECE, NLL on both splits), on the exact questions Jev's own report uses.
"New sources" = transfer-v4 (MMLU, SciQ, QNLI, PAWS, Emotion, TweetEval, held-out policy rules) β never used to train
AUBIN. "Trained sources" = decision-v7. AUBIN's training used only decision-v7/train.
Per-source accuracy, AUBIN Duo vs Jev (same clean development questions)
| source | n | AUBIN Duo | Jev |
|---|---|---|---|
| composition_held_and_or | 32 | 1.000 | 0.906 |
| composition_held_conditional | 32 | 1.000 | 0.781 |
| composition_held_or_not | 32 | 1.000 | 0.969 |
| contrastive_authorization | 40 | 1.000 | 1.000 |
| contrastive_deadline | 40 | 1.000 | 0.925 |
| emotion | 80 | 0.588 | 0.588 |
| mmlu | 80 | 0.875 | 0.900 |
| paws | 80 | 0.787 | 0.787 |
| qnli | 80 | 0.912 | 0.925 |
| sciq | 80 | 0.988 | 0.988 |
| tweet_offensive | 80 | 0.738 | 0.812 |
| agnews | 80 | 0.825 | 0.812 |
| agnews_yn | 160 | 0.894 | 0.875 |
| amazon | 80 | 0.613 | 0.588 |
| banking77 | 80 | 0.800 | 0.812 |
| boolq | 80 | 0.925 | 0.925 |
| composition_atom | 16 | 1.000 | 1.000 |
| composition_conditional | 16 | 1.000 | 0.938 |
| composition_conjunction | 16 | 1.000 | 1.000 |
| composition_disjunction | 16 | 0.938 | 0.938 |
| composition_exception | 16 | 1.000 | 1.000 |
| composition_negation | 16 | 1.000 | 1.000 |
| composition_nested_and | 16 | 0.812 | 0.812 |
| composition_nested_or | 16 | 1.000 | 1.000 |
| contrastive_age_eligibility | 24 | 1.000 | 1.000 |
| contrastive_quantity_limit | 24 | 1.000 | 0.958 |
| contrastive_return_window | 24 | 0.958 | 0.708 |
| contrastive_spend_threshold | 24 | 1.000 | 1.000 |
| dbpedia14 | 80 | 0.963 | 0.950 |
| imdb | 80 | 0.912 | 0.900 |
| mnli | 80 | 0.925 | 0.900 |
| sst5 | 80 | 0.600 | 0.637 |
| trec | 80 | 0.925 | 0.912 |
| yelp | 80 | 0.662 | 0.662 |
| yelp_yn | 80 | 0.912 | 0.863 |
AUBIN Duo ahead on 15 sources, tied on 15, behind on 5 (of 35).
Where AUBIN wins by multiples (same questions)
The costly failure of a decision model is being confidently wrong or misapplying a policy rule (refund windows, authorization, deadlines, nested conditions).
| model | split | wrong while β₯90% confident β (Jev β AUBIN) | fewer | rule/policy errors β (Jev β AUBIN) | fewer |
|---|---|---|---|---|---|
| AUBIN-12B full precision (think) | new sources | 3.66% β 1.68% | 2.2Γ | 14 β 1 / 176 | 14.0Γ |
| AUBIN-12B full precision (think) | trained sources | 5.22% β 1.34% | 3.9Γ | 13 β 0 / 224 | β (zero) |
| AUBIN-31B (fast, one pass) | new sources | 3.66% β 1.07% | 3.4Γ | 14 β 7 / 176 | 2.0Γ |
| AUBIN-31B (fast, one pass) | trained sources | 5.22% β 0.24% | 21.8Γ | 13 β 10 / 224 | 1.3Γ |
| AUBIN-31B (think) | new sources | 3.66% β 3.20% | 1.1Γ | 14 β 0 / 176 | β (zero) |
| AUBIN-31B (think) | trained sources | 5.22% β 2.22% | 2.4Γ | 13 β 6 / 224 | 2.2Γ |
| AUBIN Duo (31B + 12B, think) | new sources | 3.66% β 3.05% | 1.2Γ | 14 β 0 / 176 | β (zero) |
| AUBIN Duo (31B + 12B, think) | trained sources | 5.22% β 2.37% | 2.2Γ | 13 β 5 / 224 | 2.6Γ |
| AUBIN Duo (31B + 12B full precision, think) | new sources | 3.66% β 2.29% | 1.6Γ | 14 β 0 / 176 | β (zero) |
| AUBIN Duo (31B + 12B full precision, think) | trained sources | 5.22% β 1.98% | 2.6Γ | 13 β 5 / 224 | 2.6Γ |
| AUBIN Duo-SC (31B self-consistent think + 12B full precision) | new sources | 3.66% β 2.29% | 1.6Γ | 14 β 0 / 176 | β (zero) |
| AUBIN Duo-SC (31B self-consistent think + 12B full precision) | trained sources | 5.22% β 1.98% | 2.6Γ | 13 β 5 / 224 | 2.6Γ |
| AUBIN Trio (31B-SC + 12B full precision + Phi-4) | new sources | 3.66% β 2.74% | 1.3Γ | 14 β 0 / 176 | β (zero) |
| AUBIN Trio (31B-SC + 12B full precision + Phi-4) | trained sources | 5.22% β 2.69% | 1.9Γ | 13 β 4 / 224 | 3.2Γ |
| AUBIN Quad (31B-SC + 12B full precision + Phi-4 + Mistral-Small) | new sources | 3.66% β 2.29% | 1.6Γ | 14 β 0 / 176 | β (zero) |
| AUBIN Quad (31B-SC + 12B full precision + Phi-4 + Mistral-Small) | trained sources | 5.22% β 2.14% | 2.4Γ | 13 β 5 / 224 | 2.6Γ |
"Wrong while β₯90% confident" = share of all questions answered wrongly with confidence β₯ 0.9 (Jev: confident_error_rate from its report). AUBIN is more cautious β it is β₯90% sure on fewer questions β but when it is, it is right more often, and its overall calibration (ECE) is better. Rule/policy = composition_* + contrastive_* sources.
Against the Kev family (Kev README, development / test, clean)
On the locked test split of sources AUBIN never trained on, AUBIN-31B (fast, one pass) is the most accurate model in the table (0.898 vs Kev-27B 0.896). On Kev's own trained sources the Kev models remain ahead.
| model | new sources acc (dev / test) | trained sources acc (dev / test) | new sources Brier (dev / test) |
|---|---|---|---|
| AUBIN-31B (think) | 0.863 / 0.890 | 0.857 / 0.846 | 0.226 / 0.180 |
| AUBIN Duo-SC (31B self-consistent think + 12B full precision) | 0.863 / 0.895 | 0.862 / 0.843 | 0.210 / 0.168 |
| AUBIN Trio (31B-SC + 12B full precision + Phi-4) | 0.872 / 0.889 | 0.851 / 0.851 | 0.209 / 0.169 |
| AUBIN Quad (31B-SC + 12B full precision + Phi-4 + Mistral-Small) | 0.869 / 0.886 | 0.856 / β | 0.207 / 0.168 |
| AUBIN Duo (31B + 12B, think) | 0.867 / 0.893 | 0.859 / 0.850 | 0.214 / 0.169 |
| Kev-0.8B | 0.648 / 0.697 | 0.827 / 0.838 | 0.481 / 0.416 |
| Kev-4B | 0.817 / 0.838 | 0.873 / 0.865 | 0.269 / 0.242 |
| Kev-9B | 0.822 / 0.852 | 0.872 / 0.874 | 0.286 / 0.237 |
| Kev-27B | 0.848 / 0.896 | 0.866 / 0.870 | 0.236 / 0.164 |
| Jev (hosted) | 0.857 / β | 0.845 / β | 0.211 / β |
Computer use: Mind2Web web-agent benchmark
Each step: task + previous actions + the official Mind2Web ranker's top candidates β AUBIN picks the element, the operation (click / type / select) and the value. Metrics as in the Mind2Web paper (Deng et al., 2023): element accuracy, operation F1, step success rate. Reference rows are the paper's Table 2; GPT-4 there chose from the top 10 candidates on a 50-task subset, AUBIN from the top 50.
| split | model | element acc | op F1 | step SR |
|---|---|---|---|---|
| Cross-Task | AUBIN-31B, zero-shot (paper protocol: 5 + None) (200 steps) | 48.0 | 81.0 | 40.0 |
| Cross-Task | AUBIN-12B + Mind2Web fine-tune (paper protocol: 5 + None) (300 steps) | 40.7 | 82.1 | 36.3 |
| Cross-Task | AUBIN-12B, zero-shot (paper protocol: 5 + None) (300 steps) | 35.3 | 74.3 | 28.0 |
| Cross-Task | AUBIN-12B, zero-shot (single 50-way choice) (300 steps) | 38.3 | 77.8 | 31.3 |
| Cross-Task | GPT-4 (paper, top-10, 50 tasks) | 41.6 | 60.6 | 36.2 |
| Cross-Task | MindAct Flan-T5-XL, fine-tuned (paper) | 55.1 | 75.7 | 52.0 |
| Cross-Task | GPT-3.5 (paper) | 20.3 | 56.6 | 17.4 |
| Cross-Website | AUBIN-31B, zero-shot (paper protocol: 5 + None) (200 steps) | 41.0 | 82.0 | 33.5 |
| Cross-Website | AUBIN-12B, zero-shot (paper protocol: 5 + None) (300 steps) | 33.3 | 76.9 | 26.0 |
| Cross-Website | AUBIN-12B, zero-shot (single 50-way choice) (300 steps) | 36.0 | 79.6 | 28.3 |
| Cross-Website | GPT-4 (paper, top-10, 50 tasks) | 35.8 | 51.1 | 30.1 |
| Cross-Website | MindAct Flan-T5-XL, fine-tuned (paper) | 42.0 | 65.2 | 38.9 |
| Cross-Website | GPT-3.5 (paper) | 19.3 | 48.8 | 16.2 |
| Cross-Domain | AUBIN-31B, zero-shot (paper protocol: 5 + None) (200 steps) | 54.0 | 86.5 | 48.5 |
| Cross-Domain | AUBIN-12B, zero-shot (paper protocol: 5 + None) (300 steps) | 42.0 | 82.9 | 39.3 |
| Cross-Domain | AUBIN-12B, zero-shot (single 50-way choice) (300 steps) | 37.3 | 83.2 | 34.0 |
| Cross-Domain | GPT-4 (paper, top-10, 50 tasks) | 37.1 | 46.5 | 26.4 |
| Cross-Domain | MindAct Flan-T5-XL, fine-tuned (paper) | 42.1 | 66.5 | 39.6 |
| Cross-Domain | GPT-3.5 (paper) | 21.6 | 52.8 | 18.6 |
Game: AUBIN plays 2048 (live replay)
AUBIN-12B with a 2-step lookahead tool: the tool plays clear moves, AUBIN decides close calls (34% of moves). Same 10 seeds for every player.
| player | mean score | min | max | reached 2048 |
|---|---|---|---|---|
| AUBIN + lookahead tool | 29,009 | 7,016 | 37,132 | 8/10 |
| lookahead tool alone | 30,076 | 12,268 | 47,436 | 8/10 |
| greedy merge rule | 2,982 | 1,464 | 3,676 | 0/10 |
| random | 933 | 272 | 1,508 | 0/10 |
Real-time control: command-following grid game
Each move: a 7x7 grid as text (walls, 5 deadly lava cells, 4 objects that block movement) and a command ("go to the red key"); the controller returns one move with a calibrated confidence. Same 100 episodes (seed 0) for every row, never seen in training (training episodes use seeds 1000+). Step accuracy = the move is on a shortest safe path (BFS ground truth). The safety shield only removes moves into visible lava/walls (AubinController.act(..., allowed=...)).
AUBIN-12B-Control rows are measured with code/control_bench.py on a Colab T4; its adapter upload follows in the next update.
| controller | success | lava deaths β | step accuracy | median ms / move (T4) |
|---|---|---|---|---|
| random | 6% | 40 | 35.3% | β |
| random + safety shield | 13% | 0 | 53.4% | β |
| greedy (ignores obstacles) | 58% | 26 | 56.8% | β |
| greedy + safety shield | 89% | 0 | 87.6% | β |
| Gemma-4-E4B base, no control training (one pass per move) | 8% | 26 | 39.2% | 329 |
| Gemma-4-E4B base, no control training + safety shield | 11% | 0 | 55.4% | 330 |
| AUBIN-12B decision adapter, no control training (one pass per move) | 57% | 27 | 62.8% | 798 |
| AUBIN-12B-Control (one pass per move) | 74% | 24 | 86.4% | 854 |
| AUBIN-12B-Control + safety shield (never steps on visible lava/walls) | 92% | 0 | 91.8% | 856 |
| AUBIN-E4B-Control (one pass per move) | 58% | 28 | 60.0% | 384 |
| AUBIN-E4B-Control + safety shield (never steps on visible lava/walls) | 95% | 0 | 93.2% | 385 |
Consistency and hallucination ("gray area" test types)
The test types follow an independent public review of Jev (Onur Tirpan, YouTube, Sep 2026: same question repeated, option order reversed, question rephrased, leading question, X vs not-X, false premise, invented quote, false alarm). AUBIN-12B was run on our own synthetic student records and invoices (Turkish + English, 40 documents); Jev's numbers are the ones shown in that review on its own Turkish items, so this is not a same-item comparison, and the first set of documents is short and simple. The hard version adds distractor fields, synonym labels, look-alike numbers (IBAN/fax), a blank signature line and a long note to every document.
| test | AUBIN-12B, 40 docs | AUBIN-12B, 40 hard docs | Jev (review, TR items) |
|---|---|---|---|
| decision changed when asked again | 0/120 | 0/120 | 1/60 |
| decision changed when options reversed | 0/120 | 1/120 | 11/60 |
| decision changed when question rephrased | 0/120 | 6/120 | 5/10 |
| followed a leading question | 0/40 | 0/40 | 2/10 |
| "X" and "not X" contradicted | 0/65 | 4/65 | 0/40 |
| accepted a false premise | 0/40 | 0/40 | 0/60 |
| accepted an invented quote | 0/40 | 0/40 | β |
| said a present field is missing | 0/40 | 0/40 | 2 in 20 docs |
| answer accuracy | 240/240 | 234/240 | β |
On the hard documents AUBIN-12B still contradicted itself on "X" vs "not X" in 4/65 pairs (Jev's review shows 0/40 on its own items); this is the open item for the next training round.
AUBIN Loop: generator + decider (aubin loop)
A generator writes candidate answers; AUBIN picks among them with calibrated probabilities; if confidence is below the target, the generator is asked again with the rejected candidates as feedback. On Gemma it is one set of weights: the generator is the same base with the AUBIN adapter switched off. Any OpenAI-compatible model (ChatGPT, vLLM, Ollama) can be the generator instead (--generator openai:<model>, key from the environment).
GSM8K test, 400 fixed problems, generator = Gemma-4-12B base, same candidates for every method:
| method | accuracy | samples / problem |
|---|---|---|
| generator alone (greedy) | 0.843 | 1 |
| majority vote of 5 samples | 0.858 | 5 |
| AUBIN picks among 5 samples | 0.900 | 5 |
| AUBIN Γ vote share | 0.905 | 5 |
| AUBIN Loop (2nd round only when unsure) | 0.925 | 6.46 |
| oracle: any of 5 samples correct (upper bound) | 0.917 | 5 |
Measured speed (free Kaggle GPUs, 4-bit, batch = 1 case)
| setup | fast path | think-when-uncertain |
|---|---|---|
| AUBIN-12B, 1x T4 | 486 ms / question (runs/speed/speed12.json) | ~6-7 s extra per thought question (batched 8), 26-32% of questions |
| AUBIN-31B, 2x T4 | 1.3-2.0 s / question | ~13-15 s extra per thought question (batched 8), ~14% of questions |
| Jev (hosted API, Kev report) | ~220 ms median | β |
T4 is a 2018-era free GPU; AUBIN latency on modern datacenter GPUs has not been measured yet and is not claimed.
Use
pip install "git+https://huggingface.co/emrevrg/AUBIN-31B#subdirectory=code"
from aubin import Aubin, AubinEnsemble
m = Aubin("emrevrg/AUBIN-31B") # base + LoRA + fitted temperature + think-when-uncertain
m.decide({"ticket": "Package arrived broken, customer wants money back"},
{"route": {"type": "choice", "instructions": "Which team?",
"criteria": {"billing": "refunds", "shipping": "delivery", "tech": "software"}},
"urgent": {"type": "noul", "instructions": "Is this urgent?"}})
# AUBIN Duo-SC (headline row; 12B in full precision, e.g. 2x T4):
duo = AubinEnsemble([(Aubin("emrevrg/AUBIN-31B", think_margin=4, think_mix=0.5, think_samples=3), 0.6),
(Aubin("emrevrg/AUBIN-12B", four_bit=False, device_map="auto", temperature=3.2, think_margin=4, think_mix=0.5), 0.4)],
temperature=0.8)
# AUBIN Loop: generator + decider on the same Gemma weights (or any OpenAI-compatible generator)
from aubin import AubinLoop, GemmaGenerator
from aubin.loop import last_number
loop = AubinLoop(m, GemmaGenerator(m), extract=last_number, target=0.8, mode="aubin+vote")
loop.solve("A shop sells 3 pens for $4. How much do 12 pens cost?")
CLI: aubin decide case.json --think 4 Β· aubin loop "task" --numeric Β· server: aubin serve --port 8009 (POST /decide).
Honest notes
- Every number above is produced by the scripts in
code/from per-question outputs; nothing is hand-edited. - Every ensemble row (Duo, Duo-SC, Trio, Quad) uses one fixed selection rule: member weights and final temperature minimize
Brier + ECE on Kev's calibration split (decision-v7 calibration, 300 questions;
code/select_cal.py). The development and test splits are never used for selection; think mix is fixed at 0.5. The test columns are single confirmation runs. The new-source Brier margin is small: if the rule were calibration NLL instead, Duo-SC would be behind Jev on new-source Brier and ECE (6/8), while accuracy and NLL stay ahead on both splits. - Which members to combine was our choice among about a dozen candidate combinations of the models listed here (including 12B variants trained with distillation and data weighting); only weights and temperature come from the rule.
- AUBIN Duo-SC: the 31B member thinks with 3 samples (greedy + 2 sampled, answer probabilities averaged;
think_samples=3). - AUBIN Trio adds microsoft/phi-4 (MIT, no adapter, temperature fitted on the calibration split) as a third, different-family member.
- AUBIN Quad adds mistralai/Mistral-Small-24B-Instruct-2501 (Apache-2.0, no adapter, calibration-fitted temperature 5.5) as a fourth member; its trained-source test run hit the free-GPU session limit, so that test cell is not reported ("β").
- AUBIN-12B full precision = the same adapter on the unquantized base (fp16 on 2x T4, or bf16 on one T4 through IRIS); its temperature (3.2) is fitted on the calibration split like every other model here.
- Jev is a hosted API with lower latency (~220 ms median in Kev's report) and very low price; AUBIN's advantages are open weights, local/offline/private use, zero per-call cost on your own GPU, and fine-tunability. Think-when-uncertain adds generation time on ~14% of questions.
- Base model: google/gemma-4-31B-it (Apache-2.0). This repo holds the LoRA adapter, fitted temperature and code.
AUBIN-Learn β instant self-learning (experimental, Norovox core)
AUBIN can learn from feedback without retraining. Verified cases go into an external decision memory (a write takes well under
a millisecond with the hashing embedder), and a self-calibrator (Hedge / multiplicative weights) shifts trust per source between
the model, the memory and their fusion β so where memory is not useful yet, AUBIN keeps trusting itself. Measured on Kev's public
suites; full tables, protocol and code: reports/AUBIN_LEARN.md, code/aubin/learn.py.
- Feedback stream, kev_test: all 6 AUBIN variants improve, +1.0 to +1.3 points (best 84.3 β 85.6). Separate protocol β the label is revealed after each answer β so it is not comparable to static scores (Kev-9B 87.4 static).
- Never-seen sources (transfer, memory starts empty): β0.4 to +0.1 points β it does not hurt; a few hundred feedbacks per source are not enough to help yet.
- Fast skills, static locked test: per-source classifiers learned from memory in seconds (switched on only where dev proves them) lift kev_test for all 4 measured AUBIN variants, +0.3 to +0.6 points (ensemble 84.3 β 84.8; banking77 59.5 β 65.5).
- Raw memory (kNN) on the static test: no reliable gain (β1.0 to +0.7) β AUBIN already learned these sources.
Results, 3 October 2026 (full report: reports/AUBIN_RESULTS_2026-10-03.md)
| benchmark | AUBIN | reference |
|---|---|---|
| Typed decisions (2,000 decisions) | 77.55 (AUBIN-Learn, weights fixed before test) | Laya 76.65 Β· meraGPT 76.8 Β· Jev 72.7 |
| Kev suites, Kev's training sources (kev_test) | 85.7 (AUBIN ensemble, selected on cal split) | Kev-0.8B 83.8 Β· Kev-4B 86.5 Β· Kev-9B 87.4 |
| Kev suites, transfer test | 86.5 ensemble Β· 89.0 AUBIN-31B | β |
| Mind2Web cross-domain step SR (200 steps) | 48.5 AUBIN-31B, no web training | MindAct-XL 39.6 Β· GPT-4 26.4 |
| Grid control, 100 unseen episodes | 92% success, 0 lava deaths (12B-Control + shield) | greedy rule + same shield 89% |
On Kev's own training sources AUBIN is still 1.7 points behind Kev-9B; this is stated, not hidden.
Evening update (3 Oct, all measured, details in the results report):
| benchmark | AUBIN | reference |
|---|---|---|
| ScreenSpot click accuracy (visual computer use) | 67.7 AUBIN-E4B-Screen (zero-shot 48.5) | SeeClick 53.4 Β· CogAgent 47.4 Β· UGround-7B 73.3 Β· UI-TARS-7B 89.5 |
| Mind2Web step success, cross-task / website / domain | 47.0 / 37.0 / 43.5 (12B) Β· 47.0 / 34.5 / 42.0 (E4B, β2Γ faster) | MindAct-XL 52.0 / 38.9 / 39.6 |
| ViZDoom FPS, kills per episode | 15.1 (AUBIN-12B zero-shot, 0.55 s/move) | random 1.25 Β· scripted rule 19.6 |
| Kev training sources + learned skills (kev_dev-selected) | 86.3 (pre-declared) | Kev-9B 87.4 (not yet beaten) |
New in code/: AubinLearning.acquire_skill (learns a skill, self-tests on held-out data, enables only on proven gain), learn_skill2/3.py, fuse_multi.py, screenspot_train/eval.py, fps_vizdoom.py.
- Downloads last month
- 42

