CHSM8 move tokenizer

CHSM8's move tokenizer: any common chess notation in, the exact move tokens CHSM8 was trained on out. Rust core on shakmaty, with a Python module and a WebAssembly build of the same code.

Each move is converted to canonical UCI (castling as the king's two-square move e1g1, promotion as a lowercase letter e7e8q) and packed into 15 bits of a uint16:

bits field values
0–5 from square a1 = 0 … h8 = 63
6–11 to square a1 = 0 … h8 = 63
12–14 promotion 0 none, 1 n, 2 b, 3 r, 4 q

Piece type is not stored, because the position determines it. The model input adds it back: one row per move, (kind, from, to, piece, promotion), which analyze returns together with every legal next move.

Accepted input (notations can be mixed within one game)

examples
SAN Nf3 exd5 O-O 0-0-0 OO e8=Q+ e8Q e8(Q) e8/Q exd6 e.p. ♘f3
LAN Ng1-f3 e4xd5 e7-e8=Q Ng1f3 Ke1-g1
UCI g1f3 e7e8q e7e8Q e1g1 e1h1 (king takes rook)
PGN move numbers 1. 1... 1.e4, [headers], {comments}, ; comments, (variations), $1, !?, + #, results

Every move is checked for legality, and a LAN piece letter must match the piece on its from-square.

Install

pip install "git+https://huggingface.co/LegumMagister/chsm8-tokenizer"   # Python module chsm8_tok (needs Rust: https://rustup.rs)
git clone https://huggingface.co/LegumMagister/chsm8-tokenizer && cd chsm8-tokenizer && cargo build --release   # CLI + Rust library

The browser build is prebuilt in wasm/ (chsm8_tok.js + chsm8_tok_bg.wasm, ES module, wasm-bindgen --target web). It is what the CHSM8 demo uses. Rebuild with cargo build --release --target wasm32-unknown-unknown --lib --features wasm and wasm-bindgen --target web --out-dir wasm --no-typescript target/wasm32-unknown-unknown/release/chsm8_tok.wasm.

import init, { tokenize, analyze } from './wasm/chsm8_tok.js';
await init();
tokenize("1. e4 e5 2. Nf3");                      // Uint16Array [1804, 2356, 1350]
JSON.parse(analyze("1. e4 e5")).legal.length;     // 29 legal moves for White

Use

cargo run --release -- "1. e4 e5 2. Nf3 Nc6"           # 1804 2356 1350 2745
cargo run --release -- --uci "1. e4 e5 2. Nf3 Nc6"     # e2e4 e7e5 g1f3 b8c6
cargo run --release -- --lines < games.txt             # one game per line, about 2M moves/s
# pip install .   (needs a Rust toolchain)
import chsm8_tok
chsm8_tok.tokenize("1.e2-e4 e7-e5 2.Ng1-f3")   # [1804, 2356, 1350]
chsm8_tok.to_uci("1. e4 e5 2. O-O")            # error: illegal move 3
rows, legal = chsm8_tok.analyze("1. e4 e5")    # model rows + [(uci, san, row)] for the side to move

Tests

cargo test, and python tests/test_notations.py (needs pip install . chess):

  • one game with en passant, both castles and an underpromotion written in 12 notations, plus en passant in 4 forms;
  • 2,000 random legal games (381,943 plies) each written in 7 notations by python-chess.

All give identical tokens. On the CHSM8 datasets, the tokens match the training data exactly (7,000 pre-training games, 2,000 post-training and 2,000 top-player games checked).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support