Laya HOB Draft Pick
A Laya checkpoint fine-tuned to pick one card from a Magic: The Gathering booster pack for The Hobbit (HOB) Premier Draft on MTG Arena, given the cards already drafted. Laya is a non-generative typed-decision model: it returns a probability for every card in the pack in a single forward pass, and never generates text.
The model imitates historically strong 17Lands drafters (≥ 100 games and ≥ 58% game win rate bucket). It predicts what such a player would take, not a provably optimal pick.
Benchmark
All numbers below were re-run on 2026-09-28 on an Apple M-series Mac (MPS) with
tools/benchmarks/benchmark_hob_server.py. The evaluation drafts come from the original 17Lands
CSV and are disjoint by draft_id from every draft in the fine-tuning file (train, val and test).
| Slice (external holdout) | Picks | Options | Laya HOB | Greedy GIH WR | Uniform random |
|---|---|---|---|---|---|
| Pack 1 pick 1 | 1,000 | 14 | 70.0% (700) | 58.4% (584) | 7.1% |
| Pack 2 pick 2 | 300 | 13 | 63.0% (189) | 37.3% (112) | 7.7% |
| Picks 2–14, one per draft | 1,000 | 2–13 | 73.7% (737) | 37.5% (375) | 24.3% |
| ↳ early (picks 2–3) | 157 | 66.9% | 36.9% | ||
| ↳ mid (picks 4–9) | 461 | 70.7% | 34.3% | ||
| ↳ late (picks 10–14) | 382 | 80.1% | 41.6% |
- Greedy GIH WR: always take the pack card with the highest 17Lands games-in-hand win rate, computed on all 241,727 HOB games. This is the standard "just follow the 17Lands tier list" baseline. Its stats include games from the evaluation drafts, so it is slightly optimistic.
- The picks 2–14 score is higher mainly because late packs have few options; compare slices, not the headline number.
- A previous run of the same P1P1 file on a CUDA GPU gave 70.2%; the 0.2-point gap is numeric drift between devices.
- In-distribution check: on the fine-tuning file's own
testsplit (random 1,000 picks, all pick numbers) the model scored 68.9% vs 39.9% for greedy GIH WR (2026-09-24, CUDA).
Compared with published MTG draft bots
The chart is context, not a leaderboard. Every system was evaluated on a different set, player population, split and pick distribution; nobody else has published HOB results yet.
| System | Approach | Set / data | Top-1 | Notes |
|---|---|---|---|---|
| Laya HOB (this model) | Laya (ModernBERT-large, 421M) fine-tuned, typed choice | HOB, 17Lands, strong players | 70.0% P1P1 · 73.7% p2–14 · 68.9% in-dist. all picks | External holdout by draft_id |
| puder (Czerner, 2025) | ~7M-param transformer, imitation learning | FDN, 17Lands (~2M drafts) | 71% | Author's own test data |
| mtga-draft-engine | MLP + cross-attention ensemble | MSH, 17Lands, drafters ≥ 60% WR, last 5 days | 68.2% (MLP 67.6%, transformer 67.2%, card ratings 50.0%) | Top-3 94.9% |
| Bertram et al. 2021 | Contextual preference ranking (Siamese net) | NEO, 17Lands | 68% | As reported in UrzaGPT |
| DraftFM (Ward, 2026) | 1.6M-param choice model on card features + text embeddings | 29 sets train; BRO/FDN/MSH fully held out | 56.0% mean (50.8 / 60.4 / 56.7); 68.3% in-distribution val | Day-zero: never saw the evaluated set |
| UrzaGPT (Bertram, 2025) | Llama-3-8B + LoRA | NEO, 17Lands, 1M picks train / 10k test | 66.2% (Mistral-7B: 64.3%) | Untuned 8B models make illegal picks |
| UrzaGPT paper | GPT-4o zero-shot | NEO | 43% (38% with full card text) | Random baseline 22.1% |
| Ward et al. 2020 | NNetBot (dense net) | M19, Draftsim, 21,590 test drafts | 48.67% | DraftsimBot 44.54%, BayesBot 43.35%, Random 22.15% |
Take-aways that are supported: the model clearly beats the tier-list baseline on the same rows (+11.6 points P1P1, +25.7 P2P2, +36.2 picks 2–14), and it lands in the same band as the best published set-specific pick models (≈ 66–71%). A claim of "better than X" would require running X on these exact HOB files.
Training
| Base | convaiinnovations/laya (English, ModernBERT-large encoder, 421M params) |
| Data | 17Lands draft_data_public.HOB.PremierDraft.csv, 3,000 drafts from players with ≥ 100 games and ≥ 58% win-rate bucket |
| Card text | MTGJSON HOB.json (5.3.0+20260922), compact type/cost/colors/P-T/rarity/keywords/rules text |
| Splits | By draft_id: train / val / test (test = 12,642 picks) |
| Schedule | 4 epochs, best epoch 4 (val accuracy 69.1%, val loss 0.820) |
| Lengths | max_len=1280, head_max_len=768 |
| Temperature | [1.232, 1.2, 1.2] as stored in rl_agent_config.json |
Epoch history (validation): 65.3% → 67.2% → 68.6% → 69.1%.
Input format
The model only saw one format; use it exactly.
state = {
"game": "Magic: The Gathering", "format": "Booster Draft", "event_type": "PremierDraft",
"expansion": "HOB", "set_name": "The Hobbit",
"pack_number": 1, "pick_number": 11, "pack_size": 4,
"pool": [{"name": "Bolg's Company", "count": 1}, {"name": "Gnashing of Teeth", "count": 2}], # alphabetical
"pool_size": 3,
}
questions = {
"pick": {
"type": "choice",
"instructions": (
"You are drafting a Booster Draft pod of the Magic: The Gathering set 'The Hobbit' (HOB). "
"This is pick 11 of pack 1. Given your current pool, choose which card to take from the "
"pack shown in the options."
),
"criteria": { # the cards in the pack, alphabetical, unique names
"Gollum, Silent Slinker": "Legendary Creature — Halfling Horror, {3}{B}, B, 4/3, common, Menace. Text: Menace (This creature can't be…",
"Gundabad Opportunist": "Creature — Goblin Rogue, {3}{R}, R, 4/2, common. Text: When this creature enters, exile the top card of your library. Until the end of…",
"Iron Hills": "Land, mana value 0, colorless, common. Text: This land enters tapped. {T}: Add {R} or {W}. {2}{R}{W}, {T}, Sacrifice this…",
"Ordinary Bear": "Creature — Bear, {3}{G}, G, 4/5, common",
},
}
}
Usage
import laya
from huggingface_hub import snapshot_download
path = snapshot_download("FabioCeleste/laya-mtg-draft-picks")
agent = laya.load(path, device="cuda") # or "mps" / "cpu"
out = agent.predict(state, questions)["answers"]["pick"]
out["choice"] # card name, always one of the criteria keys
out["probabilities"] # {card: probability}, sums to 1
out["confidence"] # calibrated confidence in [0, 1]
Tested with laya==0.3.20. Throughput on an M-series Mac (MPS) was ~0.3 s per pick with 4 client threads
and serialized inference.
Limitations
- HOB Premier Draft only. Other sets, formats and cards are out of distribution.
- Imitation, not optimisation. Labels are human picks; agreement is not win rate.
- Calibration and abstain thresholds have not been evaluated on the external holdout. Do not
auto-pick on low
confidencewithout your own review threshold. - The model does not know Magic rules; everything it knows about a card comes from the text you pass
in
criteria. Supply card text from a versioned source (e.g. MTGJSON). - Not affiliated with Wizards of the Coast or 17Lands.
Related models
FabioCeleste/laya-mtg-deckbuild— chooses maindeck copy counts from a drafted pool.FabioCeleste/laya-mtg-deck-evaluator— estimates per-game win probability of a 40-card deck.
Data license
Training labels come from 17Lands public datasets, licensed CC BY 4.0. Card data from MTGJSON. The base model is Apache-2.0.
Model tree for FabioCeleste/laya-mtg-draft-picks
Base model
convaiinnovations/layaPapers for FabioCeleste/laya-mtg-draft-picks
UrzaGPT: LoRA-Tuned Large Language Models for Card Selection in Collectible Card Games
Predicting Human Card Selection in Magic: The Gathering with Contextual Preference Ranking
AI solutions for drafting in Magic: the Gathering
Evaluation results
- Top-1 agreement P1P1 (14-card pack) on 17Lands HOB PremierDraft, external holdout P1P1 (1,000 drafts)self-reported0.700
- Top-1 agreement P2P2 (13-card pack) on 17Lands HOB PremierDraft, external holdout P2P2 (300 drafts)self-reported0.630
- Top-1 agreement picks 2-14 on 17Lands HOB PremierDraft, external holdout picks 2-14 (1,000 drafts)self-reported0.737