CHSM8

CHSM8 is a research chess move-prediction model. Given a game prefix, it predicts the next chess move while applying a legal-move mask at inference. It is a project-specific PyTorch model, not a standard Transformers model.

Source code, architecture, training scripts, and data pipeline: github.com/nov3o/llm_chess.

Architecture

  • Llama-style causal chess decoder: 12 layers, width 960, 15 attention heads, SwiGLU MLP (width 3,520), RoPE, and Q/K normalization.
  • Factorized chess output heads predict move kind, source square, destination square, piece, and promotion rather than a flat move vocabulary.
  • About 165M trainable parameters. The architecture retains a frozen dummy cross-attention scaffold for compatibility with later text-conditioned work.
  • Trained in bf16 with FlashAttention, fused AdamW, gradient clipping 1.0, and weight decay 0.1.

Training data and schedule

Training uses packed Lichess standard-chess game records, represented as move sequences. The final branch reads the 3p-s packed corpus without replacement in 512-position sequences; shards 77–127 were globally reshuffled before the 85% checkpoint to remove a temporal data-order artefact.

The planned one-pass budget is 78.32B chess positions. Optimisation uses a position-indexed cosine learning-rate schedule: 5% warmup, peak LR 5e-4, and minimum LR equal to 10% of the peak. A 1% held-out split is reserved for validation. The model's legal-move penalty weight is 0.5.

Checkpoints

File Step Training exposure Description
chsm8_best_pct85.pt 1,015,809 85.00% Best checkpoint from the wide-shuffled lineage; it remained stable through the prior loss-bump window.
chsm8_latest.pt 1,194,960 99.99% Latest/final checkpoint from the training run.

Both are weights-only PyTorch dictionaries containing model_state_dict, step, total_training_time, total_tokens_processed, and cfg. Load them with FactoredChessDecoder / DecoderConfig from this repository's src/model.py.

Evaluation note

Playing evaluations use 200 colour-balanced games with greedy legal-move decoding against Stockfish constrained to UCI_Elo=2000 at 0.05 seconds per move. This is a reproducible project metric, not an official FIDE Elo rating or a general-purpose chess-engine benchmark.

Limitations

CHSM8 is a research checkpoint with no search, opening book, endgame tablebase, or standard Transformers interface. It has been evaluated only under the project protocol above; use it for research and reproduce results from the linked codebase.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support