DraftZero experiment #1: one agent for all of MTG Foundations limited

One MageZero agent trained by self-play to play any Foundations (FDN) limited deck, instead of one agent per deck. Every game drew both decks at random from 28,366 top-player decks built from 17lands data.

This release has the checkpoints, the full deck pool, every game played, and the report. It's a baseline and a set of pretrained opponents for anyone training limited agents, not a strong player.

Results

Measure Result
Training 2,507 games over 34 generations, one RunPod L40S, ~34 h, ~$28
Gen 33 vs raw search (same search, no network), final eval 110/197 (55.8%, 95% CI 49–63%)
Gen 33 vs gen 10, final eval 96/196 (49.0%, 95% CI 42–56%): training plateaued around gen 10
All milestone evals vs raw search, pooled 137/238 (57.6%, 95% CI 51–64%)
Card-value agreement with 17lands (Spearman ρ, commons, gens 10+) 0.28

The network adds a modest edge over raw search. It isn't yet strong enough for its self-play card statistics to be trusted: premium removal plays at 45–46% against 58% on 17lands. The full analysis is in report.md.

Contents

Path What
checkpoints/gen33.pt.gz Final checkpoint
checkpoints/gen10.pt.gz Where strength plateaued, and the final eval's second opponent
checkpoints/gen0.pt.gz First trained checkpoint, from 96 heuristic-search bootstrap games
decks/FDN_top_player_decks.tar.gz All 31,516 decks as XMage .dck files, plus decks.jsonl (per-deck cards, colors, player win-rate bucket)
decks/decks.tsv Train/eval split, by draft: 28,366 train, 3,150 eval
decks/eval_pairs_milestone.tsv The fixed 40-game eval (20 deck pairs × both seatings) played at every milestone
decks/eval_pairs_final.tsv The 200 raw-search games of the final eval (100 pairs, 192 decks)
run/games.jsonl Every game: both decks, colors, cards drawn, winner
run/metrics.jsonl Every metric the loop logged, per generation
run/final_eval.json, run/deck_records.tsv, run/run.json Final eval, per-deck records, run configuration and provenance
dashboards/ The run's training dashboard and format dashboard (open index.html)
report.md The experiment report

Using the checkpoints

The model is MageZero's 2-layer transformer (d_model 512). Each checkpoint carries its own feature vocabulary. Playing games with them needs three things:

Needed Where
MageZero 0.1.0: upstream v0.1.0-alpha plus 7 fork commits pip install "magezero @ git+https://github.com/danieljbrooks/MageZero@bcc76de"
The action vocabulary the policy heads index into assets/vocab/FDN_SPG.tsv in draft-zero; point MZ_ACTION_VOCAB at it
The generalist XMage build: upstream XMage plus one commit that emits actions in that vocabulary danieljbrooks/mage, branch exp1-fdn-generalist (commit 5a32441c on WillWroble/mage 2f35d9f7)

Whether the checkpoints load under MageZero v0.2.0 is untested. None of this has been run end to end outside the training harness yet.

The training harness is draft-zero (MIT), tag exp1-fdn-generalist.

Data and license

Every deck, and every human reference number in the report, comes from 17lands' public FDN Premier Draft game data (public datasets), which 17lands licenses under CC BY 4.0. The pool is every deck whose player sits in the ≥60% win-rate bucket. This release is also CC BY 4.0. If you use it, credit 17lands, and this release.

Credits

Author: Daniel Brooks. Run 2026-09-22_02-46-01, report written 2026-09-23.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading