CHSM8 move tokenizer
CHSM8's move tokenizer: any common chess notation in, the exact move tokens CHSM8 was trained on out. Rust core on shakmaty, with a Python module and a WebAssembly build of the same code.
Each move is converted to canonical UCI (castling as the king's two-square move e1g1, promotion as a lowercase
letter e7e8q) and packed into 15 bits of a uint16:
| bits | field | values |
|---|---|---|
| 0–5 | from square | a1 = 0 … h8 = 63 |
| 6–11 | to square | a1 = 0 … h8 = 63 |
| 12–14 | promotion | 0 none, 1 n, 2 b, 3 r, 4 q |
Piece type is not stored, because the position determines it. The model input adds it back: one row per move,
(kind, from, to, piece, promotion), which analyze returns together with every legal next move.
Accepted input (notations can be mixed within one game)
| examples | |
|---|---|
| SAN | Nf3 exd5 O-O 0-0-0 OO e8=Q+ e8Q e8(Q) e8/Q exd6 e.p. ♘f3 |
| LAN | Ng1-f3 e4xd5 e7-e8=Q Ng1f3 Ke1-g1 |
| UCI | g1f3 e7e8q e7e8Q e1g1 e1h1 (king takes rook) |
| PGN | move numbers 1. 1... 1.e4, [headers], {comments}, ; comments, (variations), $1, !?, + #, results |
Every move is checked for legality, and a LAN piece letter must match the piece on its from-square.
Install
pip install "git+https://huggingface.co/LegumMagister/chsm8-tokenizer" # Python module chsm8_tok (needs Rust: https://rustup.rs)
git clone https://huggingface.co/LegumMagister/chsm8-tokenizer && cd chsm8-tokenizer && cargo build --release # CLI + Rust library
The browser build is prebuilt in wasm/ (chsm8_tok.js + chsm8_tok_bg.wasm, ES module, wasm-bindgen --target web).
It is what the CHSM8 demo uses. Rebuild with
cargo build --release --target wasm32-unknown-unknown --lib --features wasm and
wasm-bindgen --target web --out-dir wasm --no-typescript target/wasm32-unknown-unknown/release/chsm8_tok.wasm.
import init, { tokenize, analyze } from './wasm/chsm8_tok.js';
await init();
tokenize("1. e4 e5 2. Nf3"); // Uint16Array [1804, 2356, 1350]
JSON.parse(analyze("1. e4 e5")).legal.length; // 29 legal moves for White
Use
cargo run --release -- "1. e4 e5 2. Nf3 Nc6" # 1804 2356 1350 2745
cargo run --release -- --uci "1. e4 e5 2. Nf3 Nc6" # e2e4 e7e5 g1f3 b8c6
cargo run --release -- --lines < games.txt # one game per line, about 2M moves/s
# pip install . (needs a Rust toolchain)
import chsm8_tok
chsm8_tok.tokenize("1.e2-e4 e7-e5 2.Ng1-f3") # [1804, 2356, 1350]
chsm8_tok.to_uci("1. e4 e5 2. O-O") # error: illegal move 3
rows, legal = chsm8_tok.analyze("1. e4 e5") # model rows + [(uci, san, row)] for the side to move
Tests
cargo test, and python tests/test_notations.py (needs pip install . chess):
- one game with en passant, both castles and an underpromotion written in 12 notations, plus en passant in 4 forms;
- 2,000 random legal games (381,943 plies) each written in 7 notations by python-chess.
All give identical tokens. On the CHSM8 datasets, the tokens match the training data exactly (7,000 pre-training games, 2,000 post-training and 2,000 top-player games checked).