System One Games

Three fine-tuned Laya decision models from STEAV's game experiments, together in getSTEAV/system-one-games: Dino, Snake and Tetris. Each game has its own LoRA adapter, trained decision head and calibration. They share a pinned open-source Laya base and one inference package.

These models select an action from typed options. They consume text and engineered simulator facts; they do not see game screenshots. Tetris selects from five placements proposed by a heuristic search. Dino reads options that explicitly mark the planner's best move.

What fine-tuning changed

CID measured the unadapted Laya checkpoint and each fine-tuned candidate on exactly the same internal held-out decisions. The base used its original decision head and calibration temperatures; the candidate used its trained head, LoRA adapter and temperatures fitted on its validation split.

Decision task Held-out decisions Stock Laya Fine-tuned Gain
Dino, explicitly labeled options 668 60.9% 100.0% +39.1 percentage points
Snake, move facts 1,201 60.9% 99.8% +38.9 percentage points
Tetris, top-five placement facts 1,600 22.9% 96.1% +73.2 percentage points

Accuracy means choosing the maximum-gold-weight option. Exact metrics, split sizes, duplicate removals, original run identities and SHA-256 provenance are in evidence/comparison.json. These are retained September 24, 2026 measurements, audited against the original CID receipts and trainer. They compare the complete fine-tuning recipe; they do not isolate the contribution of LoRA, RLCD or calibration individually.

All three candidates completed CID's seven stages: data_prep -> train -> aggregate -> eval -> benchmark -> validate -> deploy. The failed full-choice Tetris experiment is excluded from this release.

Training and base model

The base is convaiinnovations/laya, an English ModernBERT-large encoder with Laya's decision head, pinned to revision 7b928d828b7b0e022f929d9bd2e44165aa270148. Its weight file SHA-256 is 891102d372688fc2a094dac56a384bc537b87c63f21f9f3dac0be2b7cbc8d86c. Base weights are fetched from the upstream repository, rather than redistributed here.

Each candidate trains a rank-16 LoRA adapter with alpha 32 and dropout 0.05 on the Wqkv and Wo modules, plus the decision head initialized from Laya. The total trained parameter count is 30,635,009. Training uses bfloat16, batch size 32, learning rate 0.0001 and seed 42. Dino uses one supervised epoch followed by one RLCD epoch; Snake and Tetris use two supervised epochs followed by one RLCD epoch. The input limits are 128, 384 and 256 tokens respectively.

The synthetic planner-generated datasets contain 10,972 Dino decisions, 12,000 Snake decisions and 16,000 Tetris decisions before exact duplicate removal. CID drops 4,299, four and zero duplicate rows respectively, then uses a seeded 80/10/10 internal row split. Temperature fitting uses validation rows only. The released test rows preserve the original internal evaluation population.

These internal row splits are not held-out-course splits. Related situations from the same simulator courses can remain across partitions. The original gameplay experiments used five additional unseen courses; those outcomes are a separate evaluation.

Original gameplay context

The published game report, sections 4.5–4.6, reports:

  • Dino: five unseen courses of 60 seconds each, zero deaths with the safety shield on; eight deaths at the default real-time setting and three with the tuned setting when the shield was off. Timing, networking and caching affect these outcomes.
  • Snake: 99.4% agreement with the planner over 4,312 moves, mean length 61.2 versus the planner's 58.6, no wall or self collisions and one starvation stop.
  • Tetris: 157–159 lines per 400 pieces across five courses, a mean of 158.2 matching the planner, with no top-out.

Those are historical first-party gameplay summaries. Full gameplay logs and the original simulator/generator code are not bundled here. The gameplay results do not provide a stock-Laya-versus-fine-tuned survival comparison; the controlled base comparison above is decision accuracy on the retained internal rows.

Use and reproduce

Download this repository and install the inference dependencies in a Python 3.12 or 3.13 environment:

python3.12 -m venv .venv-games
source .venv-games/bin/activate
pip install 'huggingface_hub>=1.0,<3'
hf download getSTEAV/system-one-games --local-dir system-one-games
cd system-one-games
pip install -r requirements.txt
python examples/quickstart.py --game snake --device cpu

From that directory, load one game through the shared package and score its included example:

from system_one_games import GameModel
import json
from pathlib import Path

model = GameModel.from_pretrained(
    "getSTEAV/system-one-games",
    game="snake",  # "dino", "snake" or "tetris"
    device="cpu",  # use "cuda" on an NVIDIA GPU
)
example = json.loads(Path("snake/example.json").read_text())
print(model.predict(example["state"], example["question"]))

Recompute both stock and fine-tuned accuracy on all three retained test sets with:

python eval/reproduce.py --game all --device cpu --out reproduction.json

Use --device cuda for GPU evaluation. The evaluator loads models serially and does not train or fit temperatures.

The loader combines the selected adapter and decision head with the pinned base, checks the released hashes and applies the stored calibration. Keep the original option order and descriptions when reproducing a test row. The retained synthetic inputs are at evidence/<game>/test.jsonl; exact file hashes and expected original metrics are in evidence/comparison.json. Reproduction on different hardware or precision can produce small numerical differences and must be reported as a separate verification result.

A separate release check scored all 3,469 retained decisions with both model families on CPU float32. Dino and Snake's fine-tuned accuracy matched the historical measurements; Tetris scored 1,537/1,600, one fewer correct answer than the original GPU run. Stock Laya also showed small precision differences. The complete, separately labeled results are in evidence/cpu-verification.json.

The adapters require the supplied head and inference package; loading an adapter with PEFT alone does not reproduce the typed-decision contract.

Limits, data and license

Dino's options contain an explicit “Best.” label. Its perfect accuracy demonstrates learning to read the planner's labels, rather than discovering game strategy. Snake and Tetris receive planner-derived safety, distance, reachability or placement features. Their accuracy remains nearly unchanged when state text is shuffled between rows: 100.0%, 99.83% and 95.38% for Dino, Snake and Tetris. This is evidence that option facts drive the decisions, rather than understanding the board from its state description.

Use these models for research into typed decisions, simulator-specific planner imitation and adaptation to your own option schemas. Generalizing to other games, changed descriptions, other simulators or real-world decisions requires new evaluation. The original CID receipts marked the synthetic evaluation as supplemental evidence.

The upstream Laya checkpoint uses Apache-2.0. This release contains STEAV's fine-tuned parameters, new inference code and retained synthetic decision rows under the repository license. It does not redistribute Chrome assets, other game artwork, commercial game code or third-party benchmark implementations. Dino was generated using a harness derived from the Contrastive Language Models T-Rex example; Snake and Tetris rows came from the simulator/planner experiments described in the blog. The original generator implementations are not part of this package.

The models do not generate natural-language answers or implement a complete game engine. A host application supplies the typed state and options, then executes the chosen action.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for getSTEAV/system-one-games

Adapter
(25)
this model