β οΈ Validation status:
vision_adapter_clm8b.ptanddagger_prose_final_head.ptare the validated artifacts (0.939 offline ranking accuracy on Flappy Bird, real Qwen3-8B/CLIP features).tetris_head.ptis EXPERIMENTAL / NOT VALIDATED β trained on a partially-broken encoder run, val_top1 ~1.6%, uniform action probabilities (learned nothing), live play score 0. Kept for the port only; do not trust its numbers.dagger_prose_final_head.pt's live score curve did not lift off 0 in this window either β read the README's honest-results section before relying on any live-play claim.
Lev β Vision/System One experiment heads (hackathon artifacts)
Trained decision-model heads from a vision-CLM pipeline: frozen encoder + small projection heads trained with bidirectional InfoNCE, ranking a closed action set for simple 2D games. These are heads/adapters only β they require the frozen encoder + heads to run; see "How to load".
Artifacts
| file | what | size | note |
|---|---|---|---|
dagger_prose_final_head.pt |
DAgger-trained CLM-format head (state+action), prose states, Qwen3-8B encoder | 151 MB | best prose-path checkpoint |
vision_adapter_clm8b.pt |
vision adapter mapping CLIP frame feats into CLM-8B's 4096-d space | 39 MB | released CLM-8B heads stay frozen |
tetris_head.pt |
CLM-format head for the Tetris game | 63 MB | from-scratch head |
vision_adapter_report.json / tetris_report.json |
measured metrics | β | real-vs-mock labelled |
dagger_lev_flappy_head_v2.pt |
v2 Flappy head (state+action), aligned Lev physics text-state, Qwen3-8B encoder | 151 MB | best of 3 DAgger-lev candidates, verified 16-seed re-eval |
dagger_lev_flappy_head_v2_NOTES.md |
v2 eval table + training recipe | β | full per-seed scores + recipe |
Key measured results (server, AMD MI50 gfx906)
- Vision adapter ranking accuracy: 0.894 (subset 800) β 0.939 (subset 2800), vs 0.600 untrained-random adapter and 0.500 prose baseline.
- Live play: offline ranking accuracy β online control β trained adapter scored 0 pipes with 84% expert-agreement (covariate shift). DAgger heads are the on-policy correction pass.
- llm_baseline control: a small open VLM on the same frames was 0% in-schema / ~27 s/decision β System One heads answer in ms with closed-set outputs.
v2 β dagger_lev_flappy_head_v2.pt (Flappy, aligned Lev physics)
Promoted 2026-10-11 by verified re-evaluation of the three currently-trained
DAgger-lev Flappy heads on seeds 0β15 (600-decision cap,
DECISION_EVERY=3, aligned Lev physics engine, frozen Qwen3-8B last-token
encoder, closed-set head scoring). Winner rule: highest mean (tie β higher
median; within Β±0.5 mean prefer fewer zeros) β the winner took mean by 5.4
over the runner-up.
Honest comparison (live Flappy play):
| artifact | live score (pipes) | status |
|---|---|---|
v1 text_state_flappy_head.pt |
2.5 mean / 5 max | shipped baseline |
v2 dagger_lev_flappy_head_v2.pt |
11.125 mean / 8.5 median / 38 max / 0 min (1/16 zero-score seeds) | best currently-trained head; this release |
| lookahead-expert ceiling (teacher) | 43.1 mean | cap-limited at the 60 s data cap (not skill-limited); not a released head |
vision adapter vision_adapter_clm8b.pt |
0.125 live | offline ranking β online control β 0.939 offline ranking accuracy did not transfer to play (covariate shift); do not read offline metrics as a play claim |
tetris_head.pt |
0 (live) | EXPERIMENTAL / NOT VALIDATED β see the warning at the top of this README |
v2 is trained from lookahead-expert seed data + DAgger on-policy rounds on the
aligned Lev physics engine (frozen Qwen3-8B encoder, closed-set CE with
label-indexed prototypes, best-epoch selection on macro recall). The full
3-candidate eval table (with per-seed scores) and the training recipe are in
dagger_lev_flappy_head_v2_NOTES.md. Honest headline: v2 is ~4.5x the v1 mean
and the best of the trained heads, but still far below the lookahead-expert
ceiling β DAgger rounds are not monotone, and 16-seed play means carry real
seed noise (this same file measured 6.25 mean / 20 max on seeds 500-519 in its
training window).
How to load
Heads use the Contrastive-LM checkpoint format (state_head, action_head,
logit_scale, cfg). Load with torch, then score candidates by
exp(logit_scale) * cos(state_head(s), action_head(a)) over your closed
candidate set. Encoder: the prose head expects Qwen3-8B last-token embeddings
(4096-d); the vision adapter expects CLIP ViT features in β CLM space out.
See the hackathon repo for contracts.save_checkpoint / load_checkpoint.
Same-environment v1 vs v2 (added later)
Measured on the SAME aligned Lev physics (seeds 0-15, DECISION_EVERY=3, Qwen3-8B last-token embeds): v1 = 5.44 mean / 18 max, v2 = 11.06 mean / 38 max (~2x). The 2.5 figure quoted for v1 elsewhere was the old gymnasium env with per-frame decisions; v2's 11.125 vs v1's 2.5 across envs is NOT the honest comparison. Also note: old text-state heads use a different state-text coordinate convention than the aligned engine (screen-y vs height-above-floor); feed each head its training-time template or it scores 0.