Pac-Man Decision Tiny · 17,601 parameters
A small, newly trained PyTorch policy that chooses among legal Pac-Man directions. It learns from a five-second simulation/search teacher and runs locally on CPU. The safe tensor weight file is about 72 KB.
Interactive experiment explorer · Training data · Qwen LoRA comparison
What the first experiment found
Same 4,096 training examples and 512 validation examples for both learned policies:
| Policy | Validation top-choice agreement with search teacher |
|---|---|
| Original Qwen3-0.6B, option-token scoring | 25.39% |
| Qwen3-0.6B, one-epoch LoRA pilot | 32.23% |
| Uniform random legal choice (expected) | 31.67% |
| Always first listed option | 38.87% |
| This structured policy | 50.20% |
The Qwen pilot remains near chance and its 32 inspected decisions always selected the second option. This tiny baseline was added after that diagnosis, before its own gameplay test. It is a separate supervised model trained from scratch, not a Qwen adapter, a Jev model, or a multimodal model.
Independent gameplay pilot
Three held-out seeds (100, 101, 102), one fixed maze, speed 1, realtime clock, 180 simulated seconds maximum per game. Actual request latency advances the game. Training was stopped during inference. All tests ran on the same DGX host; the tiny policy runs on CPU, while Qwen runs on the GPU.
| Policy | Mean pellets / game | Mean per-game request p50 |
|---|---|---|
| base | 0.0 | 21 ms |
| greedy | 173.3 | 0 ms |
| random | 33.0 | 0 ms |
| structured | 304.3 | 1 ms |
| trained | 49.0 | 21 ms |
Zero ms denotes integer-ms reporting resolution, not zero computation.
Every per-seed score is in gameplay_results.json. This is a three-game exploratory
comparison, not a claim about unseen maps or games. The tiny policy was selected
using validation cross-entropy, never gameplay test scores. Jev was not measured.
Architecture and training
Each legal option is encoded by 15 numeric facts: current lives, remaining pellets, distance to junction, frightened time, ghost distances and presence, approaching ghost flag, edible-ghost distance, nearby food, power-pellet distance, and whether the option means turning back. Feature scales are fixed constants in the source.
A shared 15 → 64 → 64 option network is followed by mean/max set pooling and a 192 → 64 → 1 scoring network. Softmax is restricted to legal choices. Shared weights and set pooling make it permutation-equivariant: reordering options reorders the scores. The measured permutation check differed by at most 2.1e-07.
Training: NVIDIA GB10, FP32, seed 42, AdamW 0.001, weight decay 0.0001, batch 128, 100 epochs. Check validation every 10 epochs; epoch 90 has the lowest cross-entropy (1.0931) and is released. Training plus validation took 8.6 seconds, excluding data generation and model setup.
The label is a soft distribution over a search oracle's values (temperature 50). No future state or oracle score is supplied at inference. See the data card for the collection recipe, seed ranges, duplicate removal and exact file hashes.
Run it on your computer
Download structured_policy.py, predict_tiny.py, and example.json from this repo,
then run:
pip install torch safetensors huggingface_hub
python predict_tiny.py --request example.json
The script downloads this repository's tiny weight file and prints a legal direction
and candidate distribution. It executes on CPU. The model uses custom PyTorch code;
it is not an AutoModelForCausalLM checkpoint.
To reproduce training on a CUDA GPU:
python download_data.py
python train_structured.py
Scope and attribution
The policy uses structured text-derived game facts, not pixels. It has not been trained or evaluated for general language reasoning, real-world inspection or multimodal perception. Candidate scores are relative preferences, not calibrated probabilities of survival. Three seeds are a small pilot; wider evaluation remains necessary before drawing a general performance conclusion.
Game engine, feature encoder and search oracle: grapeot/decision-pacman,
commit 593ed1f59f4f97ad0bda7287ff304e9667360089, MIT (LICENSE.upstream).
This project supplies the fresh data collection, shared-option policy architecture,
new GPU training run and artifacts. Weights are Apache-2.0. Personal research project.