AUBIN by Norovox

Open, calibrated decision models that see, act and learn.

Typed decisions · computer use · real-time control · self-learning, open weights on Gemma 4.

Omni · 31B · 12B · E4B · Screen · Web · Control · Support ❤

♥ If AUBIN impresses you, press Like at the top of this page. It is the single biggest help for an independent open model.

🎬 AUBIN in 80 seconds (sound on) · Türkçe izle · vertical reel: EN / TR
Music: “Aphelion” by Scott Buckley, released under CC BY 4.0 · scottbuckley.com.au

AUBIN measured results AUBIN computer use

AUBIN-12B-Web: web agent (Mind2Web)

At each step the model chooses the page element to act on, the operation (CLICK, TYPE or SELECT), and the value. It is trained on the full Mind2Web training split and continues from the earlier AUBIN-12B web adapter.

Part of the AUBIN by Norovox family: AUBIN-12B · AUBIN-31B · full results: reports/AUBIN_RESULTS_2026-10-03.md

Mind2Web (MindAct multiple-choice protocol, top-50 candidates from the official ranker, 200 steps per split)

Scores are step success rate (%).

model cross-task cross-website cross-domain
AUBIN-12B-Web 47.0 37.0 43.5
AUBIN-E4B-Web (fast) 47.0 34.5 42.0
MindAct Flan-T5-XL 52.0 38.9 39.6
GPT-4 (MindAct) 36.2 30.1 26.4

On cross-domain it is ahead of the published MindAct baselines. On the other two splits it is behind MindAct-XL. Each split here is a 200-step sample, so every number carries about ±7 points.

Element accuracy and operation F1 per split are in reports/mind2web_final.json and in the run files.

Use

The code is code/mind2web_eval.py in the AUBIN-12B repository. Its tournament selection uses groups of 5 plus a "none" option.

python mind2web_eval.py --adapter hf:emrevrg/AUBIN-12B-Web --split test_domain --steps 200
AUBIN-Learn and the latest family results

AUBIN-Learn — instant self-learning (experimental, Norovox core)

AUBIN can learn from feedback without retraining. Verified cases go into an external decision memory (a write takes well under a millisecond with the hashing embedder), and a self-calibrator (Hedge / multiplicative weights) shifts trust per source between the model, the memory and their fusion — so where memory is not useful yet, AUBIN keeps trusting itself. Measured on Kev's public suites; full tables, protocol and code: reports/AUBIN_LEARN.md, code/aubin/learn.py.

  • Feedback stream, kev_test: all 6 AUBIN variants improve, +1.0 to +1.3 points (best 84.3 → 85.6). Separate protocol — the label is revealed after each answer — so it is not comparable to static scores (Kev-9B 87.4 static).
  • Never-seen sources (transfer, memory starts empty): −0.4 to +0.1 points — it does not hurt; a few hundred feedbacks per source are not enough to help yet.
  • Fast skills, static locked test: per-source classifiers learned from memory in seconds (switched on only where dev proves them) lift kev_test for all 4 measured AUBIN variants, +0.3 to +0.6 points (ensemble 84.3 → 84.8; banking77 59.5 → 65.5).
  • Raw memory (kNN) on the static test: no reliable gain (−1.0 to +0.7) — AUBIN already learned these sources.

Results, 3 October 2026 (full report: reports/AUBIN_RESULTS_2026-10-03.md)

benchmark AUBIN reference
Typed decisions (2,000 decisions) 77.55 (AUBIN-Learn, weights fixed before test) Laya 76.65 · meraGPT 76.8 · Jev 72.7
Kev suites, Kev's training sources (kev_test) 85.7 (AUBIN ensemble, selected on cal split) Kev-0.8B 83.8 · Kev-4B 86.5 · Kev-9B 87.4
Kev suites, transfer test 86.5 ensemble · 89.0 AUBIN-31B –
Mind2Web cross-domain step SR (200 steps) 48.5 AUBIN-31B, no web training MindAct-XL 39.6 · GPT-4 26.4
Grid control, 100 unseen episodes 92% success, 0 lava deaths (12B-Control + shield) greedy rule + same shield 89%

On Kev's own training sources AUBIN is still 1.7 points behind Kev-9B; this is stated, not hidden.

Evening update (3 Oct, all measured, details in the results report):

benchmark AUBIN reference
ScreenSpot click accuracy (visual computer use) 69.3 AUBIN-E4B-Screen (zero-shot 48.5; 67.7 after round 1) SeeClick 53.4 · CogAgent 47.4 · UGround-7B 73.3 · UI-TARS-7B 89.5
Mind2Web step success, cross-task / website / domain 47.0 / 37.0 / 43.5 (12B) · 47.0 / 34.5 / 42.0 (E4B, ≈2× faster) MindAct-XL 52.0 / 38.9 / 39.6
ViZDoom FPS, kills per episode (30 episodes) 17.3 with in-game self-learning (AUBIN-Learn) · 15.6 without (AUBIN-12B, 0.6 s/move) random 1.3 · scripted rule 18.8
Kev training sources + learned skills (selected on kev_dev + cal only) 87.08, NLL 0.43 (ensemble 85.69, NLL 0.79) Kev-9B 87.4 (not yet beaten) · Kev-27B 87.0 · Kev-4B 86.5

New in code/: AubinLearning.acquire_skill (learns a skill, self-tests on held-out data, enables only on proven gain), learn_skill2/3.py, fuse_multi.py, screenspot_train/eval.py, fps_vizdoom.py.

The AUBIN family: every ability, one interface (AubinEngine routes between them)

ability model measured
typed decisions (choice / score / yes-no), calibrated AUBIN-31B · AUBIN-12B · 12B-v3b · 12B-v3d · E4B-v3 typed-decisions 77.55 (#1) · Jev's set 8/8 · Kev unseen sources 89.8 · Kev training sources 87.08
web agent AUBIN-12B-Web · AUBIN-E4B-Web (fast) · 31B web training running Mind2Web cross-domain 43.5 (MindAct-XL 39.6)
visual computer use (click on screenshots) AUBIN-E4B-Screen · 12B screen training running ScreenSpot 69.3
real-time control AUBIN-12B-Control · AUBIN-E4B-Control 92% success, 0 lava deaths
FPS play, self-learning in game AUBIN-12B + AUBIN-Learn ViZDoom 17.3 kills/episode vs 15.6 without learning (30 episodes)
self-learning, self-acquired skills AubinLearning (learn, acquire_skill) in every repo skills switch on only after a held-out self-test proves a gain

🎬 AUBIN filmi, Türkçe (80 sn, sesi aç) · English · Müzik: “Aphelion”, Scott Buckley, CC BY 4.0

♥ Help AUBIN get seen

On Hugging Face, likes decide what people discover. Big labs have marketing teams; AUBIN has one 17-year-old student and measured results. If AUBIN is useful, interesting or just impressive to you, press ♥ Like at the top of this page and on AUBIN-Omni, AUBIN-31B and AUBIN-12B, then share it with one person who builds with AI. Every like helps an independent, open, honestly measured model get discovered.

🇹🇷 Beğenin, AUBIN'in görünür olmasını sağlar: sayfanın üstündeki ♥ Like'a basarak destek ol ve bir arkadaşına gönder.

Support Norovox

Built by a 17-year-old high school student: no sponsor, no budget, just free GPUs and AI subscriptions paid for with difficulty. Support goes into GPU compute, training, and the AI development tools this work depends on (such as Claude); supporters are credited and get early access. zgremre@gmail.com · emrevrgdev@gmail.com · Why and how →

🇹🇷 17 yaşında bir lise öğrencisinin eseri: destekçisiz, bütçesiz; ücretsiz GPU'lar ve zorlukla ödenen yapay zekâ abonelikleriyle. Desteğin GPU'ya, eğitime ve bu işin dayandığı yapay zekâ geliştirme araçlarına (Claude gibi) gider.

Built with Claude Opus 5.5 and GPT-5.6 Sol; because of OpenAI usage limits, the final stretch was completed with Claude Opus 5.5. License: Apache-2.0 (adapters and code). Base models: Google Gemma 4 (Apache-2.0).

Downloads last month
34
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for emrevrg/AUBIN-12B-Web

Adapter
(108)
this model

Collection including emrevrg/AUBIN-12B-Web