fastelo: a chess player that plays at a chosen rating, and rates you from your own moves

Two small transformers (about 7M parameters each) for human-like chess at a chosen strength:

File What it is Parameters
player.safetensors Player (policy network): the move distribution of a human of a given Elo (input: position, own Elo, opponent Elo) 6.93M
strong_player.safetensors Strong player (policy network): the player before the human phase, distilled from Leela Chess Zero; mixed in for targets above about 2100 6.93M
calibration.json Maps a target Elo to the settings that play at that strength
rating_response.json Bias correction of the rating read from the player, with its systematic error

The same player does both jobs: it plays at a target rating, and the likelihood of your moves under it, as a function of the rating, tells how strong you are. Both networks are supervised policy networks; nothing here is trained with reinforcement learning.

Related: synthetic games with phase-dependent strength (CC0), generated with these models.

How it was trained

  1. Distillation. A student transformer (64 square tokens + 3 context tokens, d=256, 8 layers) learned the policy and win/draw/loss output of the Leela Chess Zero network T1-256x10-distilled-swa-2432500 on 300M positions taken from real Lichess games (never engine self-play, to keep the variety of human positions).
  2. Human phase. Starting from that student, the player was trained on 500M positions from rated rapid and classical Lichess games (August to November 2025; 42.7M games after filtering out bots, provisional ratings and abandoned games), conditioned on the mover's and the opponent's rating.
  3. Calibration (no training). On 30,000 held-out positions analysed with Stockfish, the temperature, the share of the strong player and a small correction of the Elo input were fitted so that the policy's expected-score loss per move equals that of humans in each rating band.

Measured results

Playing strength. Top-1 match with the human move on held-out games: 53.9%. For each rating band, the Elo input under which the human moves are most likely tracks the band (correlation 0.998). Raw strength rises with the Elo input up to about 2400 and then plateaus; with the calibration the policy matches human move quality from 600 to about 2400, and comes within about one standard error at 2400-2600.

Calibrated settings (target Elo -> Elo input, temperature, share of the strong player):

Target Elo input Temperature Strong share
700 696 0.93 0.00
900 919 0.93 0.00
1100 1084 0.93 0.00
1300 1286 0.93 0.00
1500 1517 0.93 0.00
1700 1639 0.89 0.00
1900 2017 0.75 0.00
2100 2032 0.61 0.00
2300 2300 0.47 0.10

Rating a player from their own moves

The player gives p(move | position, rating). For the moves one person played, the summed log-probability as a function of the rating input is a measurement of that person's rating that uses nothing but their own decisions. fastelo.play.Rater scores every move under 12 rating inputs (400 to 2600), caps what a single move can contribute (so one slip does not decide the result) and takes the mean and width of the resulting curve. Because each move is scored on its own, the curve also shows where in a game someone played above or below their level.

rating_response.json then removes the remaining bias, measured on held-out human games with the same number of players in every 200-Elo band: for each number of moves, a monotone map f with mean f(estimate) = true rating at every true rating. Rater.rating() returns

  • value: the corrected estimate;
  • statistical error (asymmetric): the estimate plus and minus its width, mapped through f;
  • systematic error: the bias left after correction on an independent set of games. It is common to all games of a player, so playing more games does not reduce it.

Closure test on independent human games after 20 own moves (Elo):

True rating Bias, raw Bias, corrected Stat. error Inside 1 sigma Syst. Mean error of 5 games combined
400-600 +93 +28 βˆ’137 / +127 57% 32 59
600-800 +78 +11 βˆ’168 / +149 58% 19 79
800-1000 +55 -56 βˆ’170 / +218 49% 59 100
1000-1200 +114 +15 βˆ’258 / +302 56% 28 138
1200-1400 +32 -38 βˆ’303 / +316 61% 45 148
1400-1600 +35 +34 βˆ’360 / +300 69% 40 125
1600-1800 -14 +15 βˆ’326 / +270 68% 22 105
1800-2000 -62 -35 βˆ’289 / +289 70% 39 91
2000-2200 -37 +2 βˆ’263 / +300 66% 17 91
2200-2400 -91 -21 βˆ’277 / +296 76% 27 105
2400-2600 -181 -59 βˆ’301 / +232 81% 61 108

When several games are combined, each is weighted by the typical error of a game of its length, not by its own error bar: weights that depend on the estimate would bias the average.

Why there is no separate rating network

A causal transformer that reads the whole game and predicts both ratings was trained first. On human games it looked good (about 190 Elo of error after a game), but most of what it knows is the level of the game: human opponents are matched, so the opponent's play gives the rating away. Against a bot whose strength follows the estimate that is circular, and the estimate drifts. The rating from the player's own moves does not depend on the opponent:

Simulated user (20 moves) Whole-game network, bot at 700 / 1500 / 2300 Own moves, bot at 700 / 1500 / 2300
engine's best moves 1245 / 1631 / 2265 2222 / 2379 / 2412
2200 1178 / 1710 / 2182 2144 / 2264 / 2271
1500 1011 / 1496 / 1568 1636 / 1499 / 1464
900 949 / 1139 / 1244 1085 / 1232 / 1020

On real players the own-moves rating is also at least as accurate, without using the opponent at all. The whole-game network is therefore not released.

Limitations

  • Strength above about 2400-2500 is not reachable; near the top the policy is mostly the strong player played close to its best move, which is less human-like.
  • One game is a weak measurement: roughly +-300 Elo after 20 of your moves, about +-200 after 40. Several games are needed for a tight rating.
  • Ratings should only be quoted between 600 and 2200 (fastelo.response.REPORT_RANGE); beyond that, report "below 600" or "2200+". The player is calibrated from 600, and bot-generated play near the top is rated 100 to 200 Elo too high.
  • The rating was measured on rapid games; untimed or blitz play is not calibrated.
  • Ratings are on the Lichess rapid scale, not FIDE.

Use

from huggingface_hub import snapshot_download
from fastelo.play import Player, Rater   # the fastelo Python package
import chess, numpy as np

root = snapshot_download("irsotarriva/fastelo")
bot = Player(f"{root}/calibration.json")
rater = Rater(bot, response=f"{root}/rating_response.json")
board, boards, moves = chess.Board(), [], []
for _ in range(40):
    move = bot.choose(board, ply=len(moves), self_elo=1500, opp_elo=1500, rng=np.random.default_rng(len(moves)))
    boards.append(board.copy(stack=False)); moves.append(move); board.push(move)
r = rater.rating(boards, moves, side=0)                            # white, after each of its moves
print(r["value"][-1], r["lo"][-1], r["hi"][-1], r["syst"][-1])     # value, 1-sigma stat. bounds, systematic

fastelo.lite has the same player and rating in plain NumPy (no PyTorch), which is what the web demo runs.

Licence and attribution

Code and weights are released under the GNU GPL v3.0 (see LICENSE).

  • Data: Lichess open database, released under CC0.
  • Teacher: Leela Chess Zero (GPL-3.0 engine; network T1-256x10-distilled-swa-2432500 from the project's contributed networks). The Leela network itself is not redistributed here; the two players were trained on its outputs.
  • Analysis for calibration: Stockfish (GPL-3.0), used as a tool.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support