TalBot
A small (~7.3M parameter) chess policy + value network that plays in the attacking style of Mikhail Tal, paired with a PUCT Monte-Carlo search so it can actually calculate.
- Strength: ~2070 Elo in a gauntlet against Stockfish (
UCI_LimitStrength, Elo anchor 2000), 20 games, 1600 MCTS simulations/move, both colours. - Style: predicts Tal's own move ~42% of the time on held-out games.
Architecture
ResNet: 18 input planes โ 128 channels ร 10 residual blocks โ
- policy head โ 4096 logits (
from*64 + to), - value head โ
tanhin [-1, 1] from the side to move.
Search: batched PUCT MCTS (policy priors + value leaves).
Training
- Distillation from Stockfish (depth 8, MultiPV 4) on ~1.9M positions drawn from Lumbra's Gigabase (1950โ1989) โ general policy + value.
- Style adaption: policy-head-only fine-tune (body and value frozen, lr 1e-4, 8 epochs) on 123k positions where Tal was to move.
| stage | held-out policy top-1 | value MSE |
|---|---|---|
| distillation | 38.0% | 0.055 |
| + Tal policy-head fine-tune | 42.1% (Tal moves) | โ |
Usage
pip install torch python-chess numpy
# net = ResNet (see src/talbot/model.py); load model.pt (state dict + channels/blocks)
# move generation via MCTS: src/talbot/mcts.py (CompactMCTS)
Or run it as a UCI engine (drop it into any GUI):
PYTHONPATH=src python -m talbot.engine --model model.pt --sims 1600 --batch 24
Limitations
- Not a top engine. The Elo figure is anchored to Stockfish's limited-strength mode and measured over a small sample (20 games); treat it as ~2000 ยฑ ~80.
- Style is a prior bias, not a reproduction of Tal's calculation. It prefers his openings and aggressive ideas, but cannot replicate his depth of sacrifice.
- The MCTS reference implementation is pure Python and is the main speed limit.
Data & licensing
- Games: public PGN collections (Lumbra's Gigabase; a curated Tal collection).
- Labels: generated with Stockfish (GPLv3) used only as an oracle; no Stockfish binary is distributed here.
- Code and weights: MIT.
- Downloads last month
- 12