Selective Attention Freezing
Evaluation checkpoints for When Can Attention Heads Be Statically Defined? by Weixian Waylon Li, Yintao Tai, Marcio Fonseca and Shay B. Cohen.
Code | Evaluation data and splits
The catalogue covers native 124M and 1B nanoGPT models, their downstream adaptations, MQAR-adapted models and matched-budget controls.
Uploads are staged: consult checkpoints.json and use only entries with status: uploaded.
Each entry records the source checkpoint hash, exported-file hash and immutable upload revision.
These are native nanoGPT checkpoints, not Transformers AutoModel checkpoints.
Install the linked code repository and load model.pt with nanogpt.eval_nanogpt_logprobs.load_model.
The code README provides PPL, downstream, MQAR and speed-evaluation commands.
Every published export passes strict model loading, tensor-preservation checks and a bitwise comparison of original and exported BF16 logits on a fixed 64-token CUDA input. Tensor precision, fixed patterns and pruning layouts are preserved. The check confirms export equivalence; it is not a rerun of the full evaluation suite. Optimiser states, training corpora, task examples and private execution paths are excluded. These evaluation files are not sufficient to resume the original optimiser trajectory.
Use task-adapted checkpoints for SST-2, BoolQ, QuALITY and MQAR; their pretraining parents are not substitutes. Keep the native recipe, matched-token, matched-time and later-intervention cohorts distinct. The 1B pretraining study has one seed; task-finetuning seeds do not constitute independent pretraining seeds. Qwen base weights and collaborator calibration artefacts are not part of this release.