AUBIN by Norovox

Open, calibrated decision models that see, act and learn.

Typed decisions · computer use · real-time control · self-learning, open weights on Gemma 4.

Omni · 31B · 12B · E4B · Screen · Web · Control · Support ❤

♥ If AUBIN impresses you, press Like at the top of this page. It is the single biggest help for an independent open model.

🎬 AUBIN in 80 seconds (sound on) · Türkçe izle · vertical reel: EN / TR
Music: “Aphelion” by Scott Buckley, released under CC BY 4.0 · scottbuckley.com.au

AUBIN measured results AUBIN computer use

AUBIN-E4B-Control — AUBIN by Norovox

Real-time control adapter for AubinController: observation (text) + command → one move with a calibrated confidence, one forward pass per move. Part of the AUBIN family (see emrevrg/AUBIN-12B).

With the same safety shield, AUBIN-E4B-Control reaches 95% success with 0 lava deaths, vs 89% for the strongest rule baseline (greedy + shield) on the same 100 unseen episodes. The same base model without this training reaches 11% with the shield.

Benchmark: command-following grid game (100 episodes, seed 0, never seen in training)

7x7 grid as text, walls, 5 deadly lava cells, 4 objects that block movement, command like "go to the red key". Step accuracy = the move is on a shortest safe path (BFS ground truth). Training used 3,000 procedurally generated episodes (seeds 1000+, 15,161 states, soft labels over all shortest moves); the test episodes are disjoint. The safety shield only removes moves into visible lava/walls/objects (AubinController.act(..., allowed=safe_moves)); the same shield is applied to the rule baselines so its effect is separated from the model's.

controller success lava deaths ↓ step accuracy median ms / move (T4, 4-bit)
random 6% 40 35.3% –
random + safety shield 13% 0 53.4% –
greedy (ignores obstacles) 58% 26 56.8% –
greedy + safety shield 89% 0 87.6% –
Gemma-4-E4B base, no control training (one pass per move) 8% 26 39.2% 329
Gemma-4-E4B base, no control training + safety shield 11% 0 55.4% 330
AUBIN-E4B-Control (one pass per move) 58% 28 60.0% 384
AUBIN-E4B-Control + safety shield (never steps on visible lava/walls) 95% 0 93.2% 385

Use

pip install "git+https://huggingface.co/emrevrg/AUBIN-12B#subdirectory=code"
from aubin import Aubin, AubinController
ctl = AubinController(Aubin("emrevrg/AUBIN-E4B-Control"),
                      actions={"up": "move one cell up (row - 1)", "down": "move one cell down (row + 1)",
                               "left": "move one cell left (column - 1)", "right": "move one cell right (column + 1)"},
                      instructions="Choose the next move that follows the command along a shortest safe path "
                                   "(never step on lava, objects block movement).")
ctl.reset()                                    # new episode
r = ctl.act(observation, command="go to the red key", allowed={"up", "left"})   # allowed = moves that are safe right now
r["action"], r["confidence"], r["latency_ms"]

Honest notes

  • Without the shield the adapter alone is not better than a greedy rule on this game (see table); its value shows when it chooses among safe moves. The shield is a standard action mask, not a planner: it never looks ahead.
  • Latency is measured on a free Colab T4 in 4-bit, batch 1. Benchmark code: control_bench.py, training: control_train.py (in the AUBIN-12B repo code/).
  • Base model: google/gemma-4-E4B-it (Apache-2.0). This repo holds only the LoRA adapter (fp16) and its measurement file.
AUBIN-Learn and the latest family results

AUBIN-Learn — instant self-learning (experimental, Norovox core)

AUBIN can learn from feedback without retraining. Verified cases go into an external decision memory (a write takes well under a millisecond with the hashing embedder), and a self-calibrator (Hedge / multiplicative weights) shifts trust per source between the model, the memory and their fusion — so where memory is not useful yet, AUBIN keeps trusting itself. Measured on Kev's public suites; full tables, protocol and code: reports/AUBIN_LEARN.md, code/aubin/learn.py.

  • Feedback stream, kev_test: all 6 AUBIN variants improve, +1.0 to +1.3 points (best 84.3 → 85.6). Separate protocol — the label is revealed after each answer — so it is not comparable to static scores (Kev-9B 87.4 static).
  • Never-seen sources (transfer, memory starts empty): −0.4 to +0.1 points — it does not hurt; a few hundred feedbacks per source are not enough to help yet.
  • Fast skills, static locked test: per-source classifiers learned from memory in seconds (switched on only where dev proves them) lift kev_test for all 4 measured AUBIN variants, +0.3 to +0.6 points (ensemble 84.3 → 84.8; banking77 59.5 → 65.5).
  • Raw memory (kNN) on the static test: no reliable gain (−1.0 to +0.7) — AUBIN already learned these sources.

Results, 3 October 2026 (full report: reports/AUBIN_RESULTS_2026-10-03.md)

benchmark AUBIN reference
Typed decisions (2,000 decisions) 77.55 (AUBIN-Learn, weights fixed before test) Laya 76.65 · meraGPT 76.8 · Jev 72.7
Kev suites, Kev's training sources (kev_test) 85.7 (AUBIN ensemble, selected on cal split) Kev-0.8B 83.8 · Kev-4B 86.5 · Kev-9B 87.4
Kev suites, transfer test 86.5 ensemble · 89.0 AUBIN-31B –
Mind2Web cross-domain step SR (200 steps) 48.5 AUBIN-31B, no web training MindAct-XL 39.6 · GPT-4 26.4
Grid control, 100 unseen episodes 92% success, 0 lava deaths (12B-Control + shield) greedy rule + same shield 89%

On Kev's own training sources AUBIN is still 1.7 points behind Kev-9B; this is stated, not hidden.

Evening update (3 Oct, all measured, details in the results report):

benchmark AUBIN reference
ScreenSpot click accuracy (visual computer use) 69.3 AUBIN-E4B-Screen (zero-shot 48.5; 67.7 after round 1) SeeClick 53.4 · CogAgent 47.4 · UGround-7B 73.3 · UI-TARS-7B 89.5
Mind2Web step success, cross-task / website / domain 47.0 / 37.0 / 43.5 (12B) · 47.0 / 34.5 / 42.0 (E4B, ≈2× faster) MindAct-XL 52.0 / 38.9 / 39.6
ViZDoom FPS, kills per episode (30 episodes) 17.3 with in-game self-learning (AUBIN-Learn) · 15.6 without (AUBIN-12B, 0.6 s/move) random 1.3 · scripted rule 18.8
Kev training sources + learned skills (selected on kev_dev + cal only) 87.08, NLL 0.43 (ensemble 85.69, NLL 0.79) Kev-9B 87.4 (not yet beaten) · Kev-27B 87.0 · Kev-4B 86.5

New in code/: AubinLearning.acquire_skill (learns a skill, self-tests on held-out data, enables only on proven gain), learn_skill2/3.py, fuse_multi.py, screenspot_train/eval.py, fps_vizdoom.py.

The AUBIN family: every ability, one interface (AubinEngine routes between them)

ability model measured
typed decisions (choice / score / yes-no), calibrated AUBIN-31B · AUBIN-12B · 12B-v3b · 12B-v3d · E4B-v3 typed-decisions 77.55 (#1) · Jev's set 8/8 · Kev unseen sources 89.8 · Kev training sources 87.08
web agent AUBIN-12B-Web · AUBIN-E4B-Web (fast) · 31B web training running Mind2Web cross-domain 43.5 (MindAct-XL 39.6)
visual computer use (click on screenshots) AUBIN-E4B-Screen · 12B screen training running ScreenSpot 69.3
real-time control AUBIN-12B-Control · AUBIN-E4B-Control 92% success, 0 lava deaths
FPS play, self-learning in game AUBIN-12B + AUBIN-Learn ViZDoom 17.3 kills/episode vs 15.6 without learning (30 episodes)
self-learning, self-acquired skills AubinLearning (learn, acquire_skill) in every repo skills switch on only after a held-out self-test proves a gain

🎬 AUBIN filmi, Türkçe (80 sn, sesi aç) · English · Müzik: “Aphelion”, Scott Buckley, CC BY 4.0

♥ Help AUBIN get seen

On Hugging Face, likes decide what people discover. Big labs have marketing teams; AUBIN has one 17-year-old student and measured results. If AUBIN is useful, interesting or just impressive to you, press ♥ Like at the top of this page and on AUBIN-Omni, AUBIN-31B and AUBIN-12B, then share it with one person who builds with AI. Every like helps an independent, open, honestly measured model get discovered.

🇹🇷 Beğenin, AUBIN'in görünür olmasını sağlar: sayfanın üstündeki ♥ Like'a basarak destek ol ve bir arkadaşına gönder.

Support Norovox

Built by a 17-year-old high school student: no sponsor, no budget, just free GPUs and AI subscriptions paid for with difficulty. Support goes into GPU compute, training, and the AI development tools this work depends on (such as Claude); supporters are credited and get early access. zgremre@gmail.com · emrevrgdev@gmail.com · Why and how →

🇹🇷 17 yaşında bir lise öğrencisinin eseri: destekçisiz, bütçesiz; ücretsiz GPU'lar ve zorlukla ödenen yapay zekâ abonelikleriyle. Desteğin GPU'ya, eğitime ve bu işin dayandığı yapay zekâ geliştirme araçlarına (Claude gibi) gider.

Built with Claude Opus 5.5 and GPT-5.6 Sol; because of OpenAI usage limits, the final stretch was completed with Claude Opus 5.5. License: Apache-2.0 (adapters and code). Base models: Google Gemma 4 (Apache-2.0).

Downloads last month
49
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for emrevrg/AUBIN-E4B-Control

Adapter
(380)
this model

Collection including emrevrg/AUBIN-E4B-Control