⚠️ Validation status: vision_adapter_clm8b.pt and dagger_prose_final_head.pt are the validated artifacts (0.939 offline ranking accuracy on Flappy Bird, real Qwen3-8B/CLIP features). tetris_head.pt is EXPERIMENTAL / NOT VALIDATED β€” trained on a partially-broken encoder run, val_top1 ~1.6%, uniform action probabilities (learned nothing), live play score 0. Kept for the port only; do not trust its numbers. dagger_prose_final_head.pt's live score curve did not lift off 0 in this window either β€” read the README's honest-results section before relying on any live-play claim.

Lev β€” Vision/System One experiment heads (hackathon artifacts)

Trained decision-model heads from a vision-CLM pipeline: frozen encoder + small projection heads trained with bidirectional InfoNCE, ranking a closed action set for simple 2D games. These are heads/adapters only β€” they require the frozen encoder + heads to run; see "How to load".

Artifacts

file what size note
dagger_prose_final_head.pt DAgger-trained CLM-format head (state+action), prose states, Qwen3-8B encoder 151 MB best prose-path checkpoint
vision_adapter_clm8b.pt vision adapter mapping CLIP frame feats into CLM-8B's 4096-d space 39 MB released CLM-8B heads stay frozen
tetris_head.pt CLM-format head for the Tetris game 63 MB from-scratch head
vision_adapter_report.json / tetris_report.json measured metrics β€” real-vs-mock labelled
dagger_lev_flappy_head_v2.pt v2 Flappy head (state+action), aligned Lev physics text-state, Qwen3-8B encoder 151 MB best of 3 DAgger-lev candidates, verified 16-seed re-eval
dagger_lev_flappy_head_v2_NOTES.md v2 eval table + training recipe β€” full per-seed scores + recipe

Key measured results (server, AMD MI50 gfx906)

  • Vision adapter ranking accuracy: 0.894 (subset 800) β†’ 0.939 (subset 2800), vs 0.600 untrained-random adapter and 0.500 prose baseline.
  • Live play: offline ranking accuracy β‰  online control β€” trained adapter scored 0 pipes with 84% expert-agreement (covariate shift). DAgger heads are the on-policy correction pass.
  • llm_baseline control: a small open VLM on the same frames was 0% in-schema / ~27 s/decision β€” System One heads answer in ms with closed-set outputs.

v2 β€” dagger_lev_flappy_head_v2.pt (Flappy, aligned Lev physics)

Promoted 2026-10-11 by verified re-evaluation of the three currently-trained DAgger-lev Flappy heads on seeds 0–15 (600-decision cap, DECISION_EVERY=3, aligned Lev physics engine, frozen Qwen3-8B last-token encoder, closed-set head scoring). Winner rule: highest mean (tie β†’ higher median; within Β±0.5 mean prefer fewer zeros) β€” the winner took mean by 5.4 over the runner-up.

Honest comparison (live Flappy play):

artifact live score (pipes) status
v1 text_state_flappy_head.pt 2.5 mean / 5 max shipped baseline
v2 dagger_lev_flappy_head_v2.pt 11.125 mean / 8.5 median / 38 max / 0 min (1/16 zero-score seeds) best currently-trained head; this release
lookahead-expert ceiling (teacher) 43.1 mean cap-limited at the 60 s data cap (not skill-limited); not a released head
vision adapter vision_adapter_clm8b.pt 0.125 live offline ranking β‰  online control β€” 0.939 offline ranking accuracy did not transfer to play (covariate shift); do not read offline metrics as a play claim
tetris_head.pt 0 (live) EXPERIMENTAL / NOT VALIDATED β€” see the warning at the top of this README

v2 is trained from lookahead-expert seed data + DAgger on-policy rounds on the aligned Lev physics engine (frozen Qwen3-8B encoder, closed-set CE with label-indexed prototypes, best-epoch selection on macro recall). The full 3-candidate eval table (with per-seed scores) and the training recipe are in dagger_lev_flappy_head_v2_NOTES.md. Honest headline: v2 is ~4.5x the v1 mean and the best of the trained heads, but still far below the lookahead-expert ceiling β€” DAgger rounds are not monotone, and 16-seed play means carry real seed noise (this same file measured 6.25 mean / 20 max on seeds 500-519 in its training window).

How to load

Heads use the Contrastive-LM checkpoint format (state_head, action_head, logit_scale, cfg). Load with torch, then score candidates by exp(logit_scale) * cos(state_head(s), action_head(a)) over your closed candidate set. Encoder: the prose head expects Qwen3-8B last-token embeddings (4096-d); the vision adapter expects CLIP ViT features in β†’ CLM space out.

See the hackathon repo for contracts.save_checkpoint / load_checkpoint.

Same-environment v1 vs v2 (added later)

Measured on the SAME aligned Lev physics (seeds 0-15, DECISION_EVERY=3, Qwen3-8B last-token embeds): v1 = 5.44 mean / 18 max, v2 = 11.06 mean / 38 max (~2x). The 2.5 figure quoted for v1 elsewhere was the old gymnasium env with per-frame decisions; v2's 11.125 vs v1's 2.5 across envs is NOT the honest comparison. Also note: old text-state heads use a different state-text coordinate convention than the aligned engine (screen-y vs height-above-floor); feed each head its training-time template or it scores 0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support