SAE-AIU: sparse autoencoder checkpoints

Checkpoints of the newly trained sparse autoencoders (SAEs) used in the paper

Rethinking SAE Evaluation: An Atomic Interpretable Unit Framework Reveals Hidden Polysemanticity and Redundancy Riteng Zhang and Xiaoqian Wang. NeurIPS 2026, Evaluations & Datasets Track.

Code (evaluation framework, training pipeline, analysis scripts): https://github.com/ritengzhang77-max/SAE-AIU

Contents

SAEs were trained with SAELens on the MLP output of one layer of three base models.

Folder Base model Hook point Input dim Checkpoints
gpt2/ gpt2 blocks.6.hook_mlp_out 768 54
opt-125m/ facebook/opt-125m blocks.6.hook_mlp_out 768 54
tiny-stories/ tiny-stories-1L-21M blocks.0.hook_mlp_out 1024 52

For each base model the grid is six architectures x three widths (4k = 4,096; 16k = 16,384; 65k = 65,536 features) x three sparsity settings:

Architecture (folder prefix) Sparsity setting
standard L1 coefficient 1.0 / 2.0 / 5.0
gated L1 coefficient 1.0 / 2.0 / 5.0
jumprelu L0 coefficient 1.0 / 3.0 / 5.0
topk k = 50 / 100 / 200
batchtopk k = 50 / 100 / 200
matryoshka_batchtopk k = 50 / 100 / 200

combinations.json lists all 162 training configurations. Two TinyStories configurations did not complete training and have no checkpoint: matryoshka_batchtopk_L0_65k_k100 and matryoshka_batchtopk_L0_65k_k200.

File layout

<base model>/checkpoints/<architecture>_L<layer>_<width>_<sparsity setting>/
    sae_weights.safetensors   # SAE parameters
    sparsity.safetensors      # feature sparsity statistics saved by SAELens
    cfg.json                  # SAE configuration
    runner_cfg.json           # training configuration

Example: gpt2/checkpoints/topk_L6_16k_k100/.

Loading

The checkpoints are in SAELens on-disk format:

from huggingface_hub import snapshot_download
from sae_lens import SAE

root = snapshot_download(
    "ritengzhang/SAE-AIU",
    allow_patterns=["gpt2/checkpoints/topk_L6_16k_k100/*"],
)
sae = SAE.load_from_disk(f"{root}/gpt2/checkpoints/topk_L6_16k_k100", device="cpu")

The code repository's benchmark_training/load_benchmark_sae.py loads a checkpoint together with its base model.

Training data

GPT-2 and OPT-125M SAEs were trained on apollo-research/Skylion007-openwebtext-tokenizer-gpt2; TinyStories SAEs on apollo-research/roneneldan-TinyStories-tokenizer-gpt2. Each run uses 18,000 training steps at 4,096 tokens per batch. Full hyperparameters are in each runner_cfg.json.

License

The SAE checkpoints in this repository are released under the MIT License. The base models are not redistributed here and remain under their own licenses (GPT-2: MIT; OPT-125M: OPT model license, non-commercial research use; TinyStories model: no license declared on its model card).

Citation

@inproceedings{zhang2026rethinking,
  title     = {Rethinking SAE Evaluation: An Atomic Interpretable Unit Framework Reveals Hidden Polysemanticity and Redundancy},
  author    = {Zhang, Riteng and Wang, Xiaoqian},
  booktitle = {Advances in Neural Information Processing Systems (Evaluations and Datasets Track)},
  year      = {2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support