DraftZero experiment #1: one agent for all of MTG Foundations limited
One MageZero agent trained by self-play to play any Foundations (FDN) limited deck, instead of one agent per deck. Every game drew both decks at random from 28,366 top-player decks built from 17lands data.
This release has the checkpoints, the full deck pool, every game played, and the report. It's a baseline and a set of pretrained opponents for anyone training limited agents, not a strong player.
Results
| Measure | Result |
|---|---|
| Training | 2,507 games over 34 generations, one RunPod L40S, ~34 h, ~$28 |
| Gen 33 vs raw search (same search, no network), final eval | 110/197 (55.8%, 95% CI 49–63%) |
| Gen 33 vs gen 10, final eval | 96/196 (49.0%, 95% CI 42–56%): training plateaued around gen 10 |
| All milestone evals vs raw search, pooled | 137/238 (57.6%, 95% CI 51–64%) |
| Card-value agreement with 17lands (Spearman ρ, commons, gens 10+) | 0.28 |
The network adds a modest edge over raw search. It isn't yet strong enough for its
self-play card statistics to be trusted: premium removal plays at 45–46% against 58% on
17lands. The full analysis is in report.md.
Contents
| Path | What |
|---|---|
checkpoints/gen33.pt.gz |
Final checkpoint |
checkpoints/gen10.pt.gz |
Where strength plateaued, and the final eval's second opponent |
checkpoints/gen0.pt.gz |
First trained checkpoint, from 96 heuristic-search bootstrap games |
decks/FDN_top_player_decks.tar.gz |
All 31,516 decks as XMage .dck files, plus decks.jsonl (per-deck cards, colors, player win-rate bucket) |
decks/decks.tsv |
Train/eval split, by draft: 28,366 train, 3,150 eval |
decks/eval_pairs_milestone.tsv |
The fixed 40-game eval (20 deck pairs × both seatings) played at every milestone |
decks/eval_pairs_final.tsv |
The 200 raw-search games of the final eval (100 pairs, 192 decks) |
run/games.jsonl |
Every game: both decks, colors, cards drawn, winner |
run/metrics.jsonl |
Every metric the loop logged, per generation |
run/final_eval.json, run/deck_records.tsv, run/run.json |
Final eval, per-deck records, run configuration and provenance |
dashboards/ |
The run's training dashboard and format dashboard (open index.html) |
report.md |
The experiment report |
Using the checkpoints
The model is MageZero's 2-layer transformer (d_model 512). Each checkpoint carries its own feature vocabulary. Playing games with them needs three things:
| Needed | Where |
|---|---|
| MageZero 0.1.0: upstream v0.1.0-alpha plus 7 fork commits | pip install "magezero @ git+https://github.com/danieljbrooks/MageZero@bcc76de" |
| The action vocabulary the policy heads index into | assets/vocab/FDN_SPG.tsv in draft-zero; point MZ_ACTION_VOCAB at it |
| The generalist XMage build: upstream XMage plus one commit that emits actions in that vocabulary | danieljbrooks/mage, branch exp1-fdn-generalist (commit 5a32441c on WillWroble/mage 2f35d9f7) |
Whether the checkpoints load under MageZero v0.2.0 is untested. None of this has been run end to end outside the training harness yet.
The training harness is draft-zero (MIT), tag
exp1-fdn-generalist.
Data and license
Every deck, and every human reference number in the report, comes from 17lands' public FDN Premier Draft game data (public datasets), which 17lands licenses under CC BY 4.0. The pool is every deck whose player sits in the ≥60% win-rate bucket. This release is also CC BY 4.0. If you use it, credit 17lands, and this release.
Credits
- 17lands for the public data, and the players who share it.
- Will Wroble for MageZero, and for advice on this run.
- chrismaghuhn for advice on compute and on performance (WillWroble/MageZero#3).
- The XMage project for the rules engine.
Author: Daniel Brooks. Run 2026-09-22_02-46-01, report written 2026-09-23.