TranSGrid Models
This repository provides the trained Transformer checkpoints for TranSGrid, introduced in What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis.
TranSGrid is a grid-transformation testbed for studying systematic generalization through deductive, inductive, and abductive reasoning. Given an initial board and a target board, a model needs to generate a sequence of actions that transforms the initial board into the target. The action space includes ten operation types on rows, columns, and local 2 × 2 blocks.
The released models operate on 6 × 6 boards containing digits from 0 to 9. Predictions are evaluated by executing the generated actions: a prediction is correct if it reaches the target board, even if it differs from the reference sequence.
Links: Paper · Code · Dataset · Interactive Demo
Available Models
Checkpoints are organized into four model configurations:
| Directory | Configuration |
|---|---|
flat_board/ |
FLAT + BOARD |
flat_pair/ |
FLAT + PAIR |
grid_board/ |
GRID + BOARD |
grid_pair/ |
GRID + PAIR |
Each configuration contains seven encoder–decoder Transformer checkpoints, named level1.pt through level7.pt. The GRID+PAIR models range from 0.96M parameters at Level 1 to 88.43M parameters at Level 7.
The examples below use grid_pair/level7.pt. To run another model, download its checkpoint and update the --ckpt argument.
Installation
The checkpoints are PyTorch .pt files. Use them with the model implementation and evaluation scripts in the GitHub repository.
Python 3.10 or newer is required. Run all subsequent commands from the code repository's root directory.
git clone https://github.com/BlueWhaleLab/TranSGrid.git
cd TranSGrid
conda create -p ./.conda python=3.12
conda activate ./.conda
pip install -r requirements.txt
pip install huggingface_hub
Download a Checkpoint
Download the GRID+PAIR Level 7 model:
hf download BlueWhaleLab/TranSGrid \
grid_pair/level7.pt \
--local-dir checkpoints
The checkpoint will be saved to checkpoints/grid_pair/level7.pt.
For example, to download the FLAT+BOARD Level 3 model:
hf download BlueWhaleLab/TranSGrid \
flat_board/level3.pt \
--local-dir checkpoints
Run Evaluation
Greedy Decoding
Evaluate the model on the TranSGrid benchmark:
python eval.py \
--ckpt checkpoints/grid_pair/level7.pt \
--data data/transgrid.jsonl \
--beams 1 \
--batch-size 128 \
--save results/TranSGrid/level7_grid_pair-greedy-bs128.jsonl \
--group-by deductive_score inductive_score abductive_level stop_reason
Beam Search (Top@8)
Generate eight candidates per instance:
python eval.py \
--ckpt checkpoints/grid_pair/level7.pt \
--data data/transgrid.jsonl \
--beams 8 \
--batch-size 128 \
--save results/TranSGrid/level7_grid_pair-beam8-bs128.jsonl \
--group-by deductive_score inductive_score abductive_level stop_reason
Top@8 counts an instance as solved if any of the eight candidates reaches the target board. Each saved JSONL row contains one selected prediction, its solve status, and the rank of the first successful candidate; it does not contain all eight candidates.
Reduce --batch-size if GPU memory is limited.
Held-out Test Set
Download the held-out Test split:
hf download BlueWhaleLab/TranSGrid \
test.jsonl \
--repo-type dataset \
--local-dir data
Then evaluate it:
python eval.py \
--ckpt checkpoints/grid_pair/level7.pt \
--data data/test.jsonl \
--beams 8 \
--batch-size 128 \
--save results/Test/level7_grid_pair-beam8-bs128.jsonl
View Saved Results
Print aggregate and grouped results without running decoding again:
python report.py results/TranSGrid \
--axes deductive_score inductive_score abductive_level stop_reason
For dataset generation, model training, and evaluation on the Decoupled and SCAN variants, see the GitHub README.
Citation
If you use these models or TranSGrid in your research, please cite:
@article{qi2026current,
title={What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis},
author={Qi, Chengwen and Ye, Deheng and Bian, Yatao},
journal={arXiv preprint arXiv:2609.19212},
year={2026}
}