Laya HOB Draft Pick

A Laya checkpoint fine-tuned to pick one card from a Magic: The Gathering booster pack for The Hobbit (HOB) Premier Draft on MTG Arena, given the cards already drafted. Laya is a non-generative typed-decision model: it returns a probability for every card in the pack in a single forward pass, and never generates text.

The model imitates historically strong 17Lands drafters (≥ 100 games and ≥ 58% game win rate bucket). It predicts what such a player would take, not a provably optimal pick.

Benchmark

All numbers below were re-run on 2026-09-28 on an Apple M-series Mac (MPS) with tools/benchmarks/benchmark_hob_server.py. The evaluation drafts come from the original 17Lands CSV and are disjoint by draft_id from every draft in the fine-tuning file (train, val and test).

Top-1 agreement on HOB external holdout: Laya 70.0% / 63.0% / 73.7% vs greedy GIH WR 58.4% / 37.3% / 37.5% vs random 7.1% / 7.7% / 24.3%

Slice (external holdout) Picks Options Laya HOB Greedy GIH WR Uniform random
Pack 1 pick 1 1,000 14 70.0% (700) 58.4% (584) 7.1%
Pack 2 pick 2 300 13 63.0% (189) 37.3% (112) 7.7%
Picks 2–14, one per draft 1,000 2–13 73.7% (737) 37.5% (375) 24.3%
↳ early (picks 2–3) 157 66.9% 36.9%
↳ mid (picks 4–9) 461 70.7% 34.3%
↳ late (picks 10–14) 382 80.1% 41.6%
  • Greedy GIH WR: always take the pack card with the highest 17Lands games-in-hand win rate, computed on all 241,727 HOB games. This is the standard "just follow the 17Lands tier list" baseline. Its stats include games from the evaluation drafts, so it is slightly optimistic.
  • The picks 2–14 score is higher mainly because late packs have few options; compare slices, not the headline number.
  • A previous run of the same P1P1 file on a CUDA GPU gave 70.2%; the 0.2-point gap is numeric drift between devices.
  • In-distribution check: on the fine-tuning file's own test split (random 1,000 picks, all pick numbers) the model scored 68.9% vs 39.9% for greedy GIH WR (2026-09-24, CUDA).

Compared with published MTG draft bots

Published top-1 human-pick agreement of MTG draft bots compared with this model

The chart is context, not a leaderboard. Every system was evaluated on a different set, player population, split and pick distribution; nobody else has published HOB results yet.

System Approach Set / data Top-1 Notes
Laya HOB (this model) Laya (ModernBERT-large, 421M) fine-tuned, typed choice HOB, 17Lands, strong players 70.0% P1P1 · 73.7% p2–14 · 68.9% in-dist. all picks External holdout by draft_id
puder (Czerner, 2025) ~7M-param transformer, imitation learning FDN, 17Lands (~2M drafts) 71% Author's own test data
mtga-draft-engine MLP + cross-attention ensemble MSH, 17Lands, drafters ≥ 60% WR, last 5 days 68.2% (MLP 67.6%, transformer 67.2%, card ratings 50.0%) Top-3 94.9%
Bertram et al. 2021 Contextual preference ranking (Siamese net) NEO, 17Lands 68% As reported in UrzaGPT
DraftFM (Ward, 2026) 1.6M-param choice model on card features + text embeddings 29 sets train; BRO/FDN/MSH fully held out 56.0% mean (50.8 / 60.4 / 56.7); 68.3% in-distribution val Day-zero: never saw the evaluated set
UrzaGPT (Bertram, 2025) Llama-3-8B + LoRA NEO, 17Lands, 1M picks train / 10k test 66.2% (Mistral-7B: 64.3%) Untuned 8B models make illegal picks
UrzaGPT paper GPT-4o zero-shot NEO 43% (38% with full card text) Random baseline 22.1%
Ward et al. 2020 NNetBot (dense net) M19, Draftsim, 21,590 test drafts 48.67% DraftsimBot 44.54%, BayesBot 43.35%, Random 22.15%

Take-aways that are supported: the model clearly beats the tier-list baseline on the same rows (+11.6 points P1P1, +25.7 P2P2, +36.2 picks 2–14), and it lands in the same band as the best published set-specific pick models (≈ 66–71%). A claim of "better than X" would require running X on these exact HOB files.

Training

Base convaiinnovations/laya (English, ModernBERT-large encoder, 421M params)
Data 17Lands draft_data_public.HOB.PremierDraft.csv, 3,000 drafts from players with ≥ 100 games and ≥ 58% win-rate bucket
Card text MTGJSON HOB.json (5.3.0+20260922), compact type/cost/colors/P-T/rarity/keywords/rules text
Splits By draft_id: train / val / test (test = 12,642 picks)
Schedule 4 epochs, best epoch 4 (val accuracy 69.1%, val loss 0.820)
Lengths max_len=1280, head_max_len=768
Temperature [1.232, 1.2, 1.2] as stored in rl_agent_config.json

Epoch history (validation): 65.3% → 67.2% → 68.6% → 69.1%.

Input format

The model only saw one format; use it exactly.

state = {
    "game": "Magic: The Gathering", "format": "Booster Draft", "event_type": "PremierDraft",
    "expansion": "HOB", "set_name": "The Hobbit",
    "pack_number": 1, "pick_number": 11, "pack_size": 4,
    "pool": [{"name": "Bolg's Company", "count": 1}, {"name": "Gnashing of Teeth", "count": 2}],  # alphabetical
    "pool_size": 3,
}
questions = {
    "pick": {
        "type": "choice",
        "instructions": (
            "You are drafting a Booster Draft pod of the Magic: The Gathering set 'The Hobbit' (HOB). "
            "This is pick 11 of pack 1. Given your current pool, choose which card to take from the "
            "pack shown in the options."
        ),
        "criteria": {  # the cards in the pack, alphabetical, unique names
            "Gollum, Silent Slinker": "Legendary Creature — Halfling Horror, {3}{B}, B, 4/3, common, Menace. Text: Menace (This creature can't be…",
            "Gundabad Opportunist": "Creature — Goblin Rogue, {3}{R}, R, 4/2, common. Text: When this creature enters, exile the top card of your library. Until the end of…",
            "Iron Hills": "Land, mana value 0, colorless, common. Text: This land enters tapped. {T}: Add {R} or {W}. {2}{R}{W}, {T}, Sacrifice this…",
            "Ordinary Bear": "Creature — Bear, {3}{G}, G, 4/5, common",
        },
    }
}

Usage

import laya
from huggingface_hub import snapshot_download

path = snapshot_download("FabioCeleste/laya-mtg-draft-picks")
agent = laya.load(path, device="cuda")  # or "mps" / "cpu"
out = agent.predict(state, questions)["answers"]["pick"]
out["choice"]         # card name, always one of the criteria keys
out["probabilities"]  # {card: probability}, sums to 1
out["confidence"]     # calibrated confidence in [0, 1]

Tested with laya==0.3.20. Throughput on an M-series Mac (MPS) was ~0.3 s per pick with 4 client threads and serialized inference.

Limitations

  • HOB Premier Draft only. Other sets, formats and cards are out of distribution.
  • Imitation, not optimisation. Labels are human picks; agreement is not win rate.
  • Calibration and abstain thresholds have not been evaluated on the external holdout. Do not auto-pick on low confidence without your own review threshold.
  • The model does not know Magic rules; everything it knows about a card comes from the text you pass in criteria. Supply card text from a versioned source (e.g. MTGJSON).
  • Not affiliated with Wizards of the Coast or 17Lands.

Related models

Data license

Training labels come from 17Lands public datasets, licensed CC BY 4.0. Card data from MTGJSON. The base model is Apache-2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.4B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FabioCeleste/laya-mtg-draft-picks

Finetuned
(94)
this model

Papers for FabioCeleste/laya-mtg-draft-picks

Evaluation results

  • Top-1 agreement P1P1 (14-card pack) on 17Lands HOB PremierDraft, external holdout P1P1 (1,000 drafts)
    self-reported
    0.700
  • Top-1 agreement P2P2 (13-card pack) on 17Lands HOB PremierDraft, external holdout P2P2 (300 drafts)
    self-reported
    0.630
  • Top-1 agreement picks 2-14 on 17Lands HOB PremierDraft, external holdout picks 2-14 (1,000 drafts)
    self-reported
    0.737