fastelo: a chess player that plays at a chosen rating, and rates you from your own moves
Two small transformers (about 7M parameters each) for human-like chess at a chosen strength:
| File | What it is | Parameters |
|---|---|---|
player.safetensors |
Player (policy network): the move distribution of a human of a given Elo (input: position, own Elo, opponent Elo) | 6.93M |
strong_player.safetensors |
Strong player (policy network): the player before the human phase, distilled from Leela Chess Zero; mixed in for targets above about 2100 | 6.93M |
calibration.json |
Maps a target Elo to the settings that play at that strength | |
rating_response.json |
Bias correction of the rating read from the player, with its systematic error |
The same player does both jobs: it plays at a target rating, and the likelihood of your moves under it, as a function of the rating, tells how strong you are. Both networks are supervised policy networks; nothing here is trained with reinforcement learning.
Related: synthetic games with phase-dependent strength (CC0), generated with these models.
How it was trained
- Distillation. A student transformer (64 square tokens + 3 context tokens, d=256, 8 layers) learned the policy and win/draw/loss output of the Leela Chess Zero network T1-256x10-distilled-swa-2432500 on 300M positions taken from real Lichess games (never engine self-play, to keep the variety of human positions).
- Human phase. Starting from that student, the player was trained on 500M positions from rated rapid and classical Lichess games (August to November 2025; 42.7M games after filtering out bots, provisional ratings and abandoned games), conditioned on the mover's and the opponent's rating.
- Calibration (no training). On 30,000 held-out positions analysed with Stockfish, the temperature, the share of the strong player and a small correction of the Elo input were fitted so that the policy's expected-score loss per move equals that of humans in each rating band.
Measured results
Playing strength. Top-1 match with the human move on held-out games: 53.9%. For each rating band, the Elo input under which the human moves are most likely tracks the band (correlation 0.998). Raw strength rises with the Elo input up to about 2400 and then plateaus; with the calibration the policy matches human move quality from 600 to about 2400, and comes within about one standard error at 2400-2600.
Calibrated settings (target Elo -> Elo input, temperature, share of the strong player):
| Target | Elo input | Temperature | Strong share |
|---|---|---|---|
| 700 | 696 | 0.93 | 0.00 |
| 900 | 919 | 0.93 | 0.00 |
| 1100 | 1084 | 0.93 | 0.00 |
| 1300 | 1286 | 0.93 | 0.00 |
| 1500 | 1517 | 0.93 | 0.00 |
| 1700 | 1639 | 0.89 | 0.00 |
| 1900 | 2017 | 0.75 | 0.00 |
| 2100 | 2032 | 0.61 | 0.00 |
| 2300 | 2300 | 0.47 | 0.10 |
Rating a player from their own moves
The player gives p(move | position, rating). For the moves one person played, the summed
log-probability as a function of the rating input is a measurement of that person's rating that
uses nothing but their own decisions. fastelo.play.Rater scores every move under 12 rating
inputs (400 to 2600), caps what a single move can contribute (so one slip does not decide the
result) and takes the mean and width of the resulting curve. Because each move is scored on its
own, the curve also shows where in a game someone played above or below their level.
rating_response.json then removes the remaining bias, measured on held-out human games with
the same number of players in every 200-Elo band: for each number of moves, a monotone map f
with mean f(estimate) = true rating at every true rating. Rater.rating() returns
- value: the corrected estimate;
- statistical error (asymmetric): the estimate plus and minus its width, mapped through f;
- systematic error: the bias left after correction on an independent set of games. It is common to all games of a player, so playing more games does not reduce it.
Closure test on independent human games after 20 own moves (Elo):
| True rating | Bias, raw | Bias, corrected | Stat. error | Inside 1 sigma | Syst. | Mean error of 5 games combined |
|---|---|---|---|---|---|---|
| 400-600 | +93 | +28 | β137 / +127 | 57% | 32 | 59 |
| 600-800 | +78 | +11 | β168 / +149 | 58% | 19 | 79 |
| 800-1000 | +55 | -56 | β170 / +218 | 49% | 59 | 100 |
| 1000-1200 | +114 | +15 | β258 / +302 | 56% | 28 | 138 |
| 1200-1400 | +32 | -38 | β303 / +316 | 61% | 45 | 148 |
| 1400-1600 | +35 | +34 | β360 / +300 | 69% | 40 | 125 |
| 1600-1800 | -14 | +15 | β326 / +270 | 68% | 22 | 105 |
| 1800-2000 | -62 | -35 | β289 / +289 | 70% | 39 | 91 |
| 2000-2200 | -37 | +2 | β263 / +300 | 66% | 17 | 91 |
| 2200-2400 | -91 | -21 | β277 / +296 | 76% | 27 | 105 |
| 2400-2600 | -181 | -59 | β301 / +232 | 81% | 61 | 108 |
When several games are combined, each is weighted by the typical error of a game of its length, not by its own error bar: weights that depend on the estimate would bias the average.
Why there is no separate rating network
A causal transformer that reads the whole game and predicts both ratings was trained first. On human games it looked good (about 190 Elo of error after a game), but most of what it knows is the level of the game: human opponents are matched, so the opponent's play gives the rating away. Against a bot whose strength follows the estimate that is circular, and the estimate drifts. The rating from the player's own moves does not depend on the opponent:
| Simulated user (20 moves) | Whole-game network, bot at 700 / 1500 / 2300 | Own moves, bot at 700 / 1500 / 2300 |
|---|---|---|
| engine's best moves | 1245 / 1631 / 2265 | 2222 / 2379 / 2412 |
| 2200 | 1178 / 1710 / 2182 | 2144 / 2264 / 2271 |
| 1500 | 1011 / 1496 / 1568 | 1636 / 1499 / 1464 |
| 900 | 949 / 1139 / 1244 | 1085 / 1232 / 1020 |
On real players the own-moves rating is also at least as accurate, without using the opponent at all. The whole-game network is therefore not released.
Limitations
- Strength above about 2400-2500 is not reachable; near the top the policy is mostly the strong player played close to its best move, which is less human-like.
- One game is a weak measurement: roughly +-300 Elo after 20 of your moves, about +-200 after 40. Several games are needed for a tight rating.
- Ratings should only be quoted between 600 and 2200 (
fastelo.response.REPORT_RANGE); beyond that, report "below 600" or "2200+". The player is calibrated from 600, and bot-generated play near the top is rated 100 to 200 Elo too high. - The rating was measured on rapid games; untimed or blitz play is not calibrated.
- Ratings are on the Lichess rapid scale, not FIDE.
Use
from huggingface_hub import snapshot_download
from fastelo.play import Player, Rater # the fastelo Python package
import chess, numpy as np
root = snapshot_download("irsotarriva/fastelo")
bot = Player(f"{root}/calibration.json")
rater = Rater(bot, response=f"{root}/rating_response.json")
board, boards, moves = chess.Board(), [], []
for _ in range(40):
move = bot.choose(board, ply=len(moves), self_elo=1500, opp_elo=1500, rng=np.random.default_rng(len(moves)))
boards.append(board.copy(stack=False)); moves.append(move); board.push(move)
r = rater.rating(boards, moves, side=0) # white, after each of its moves
print(r["value"][-1], r["lo"][-1], r["hi"][-1], r["syst"][-1]) # value, 1-sigma stat. bounds, systematic
fastelo.lite has the same player and rating in plain NumPy (no PyTorch), which is what the web
demo runs.
Licence and attribution
Code and weights are released under the GNU GPL v3.0 (see LICENSE).
- Data: Lichess open database, released under CC0.
- Teacher: Leela Chess Zero (GPL-3.0 engine; network T1-256x10-distilled-swa-2432500 from the project's contributed networks). The Leela network itself is not redistributed here; the two players were trained on its outputs.
- Analysis for calibration: Stockfish (GPL-3.0), used as a tool.