Instructions to use ritengzhang/SAE-AIU with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- SAELens
How to use ritengzhang/SAE-AIU with SAELens:
# pip install sae-lens from sae_lens import SAE sae, cfg_dict, sparsity = SAE.from_pretrained( release = "RELEASE_ID", # e.g., "gpt2-small-res-jb". See other options in https://github.com/jbloomAus/SAELens/blob/main/sae_lens/pretrained_saes.yaml sae_id = "SAE_ID", # e.g., "blocks.8.hook_resid_pre". Won't always be a hook point ) - Notebooks
- Google Colab
- Kaggle
SAE-AIU: sparse autoencoder checkpoints
Checkpoints of the newly trained sparse autoencoders (SAEs) used in the paper
Rethinking SAE Evaluation: An Atomic Interpretable Unit Framework Reveals Hidden Polysemanticity and Redundancy Riteng Zhang and Xiaoqian Wang. NeurIPS 2026, Evaluations & Datasets Track.
Code (evaluation framework, training pipeline, analysis scripts): https://github.com/ritengzhang77-max/SAE-AIU
Contents
SAEs were trained with SAELens on the MLP output of one layer of three base models.
| Folder | Base model | Hook point | Input dim | Checkpoints |
|---|---|---|---|---|
gpt2/ |
gpt2 |
blocks.6.hook_mlp_out |
768 | 54 |
opt-125m/ |
facebook/opt-125m |
blocks.6.hook_mlp_out |
768 | 54 |
tiny-stories/ |
tiny-stories-1L-21M |
blocks.0.hook_mlp_out |
1024 | 52 |
For each base model the grid is six architectures x three widths (4k = 4,096; 16k = 16,384; 65k = 65,536 features) x three sparsity settings:
| Architecture (folder prefix) | Sparsity setting |
|---|---|
standard |
L1 coefficient 1.0 / 2.0 / 5.0 |
gated |
L1 coefficient 1.0 / 2.0 / 5.0 |
jumprelu |
L0 coefficient 1.0 / 3.0 / 5.0 |
topk |
k = 50 / 100 / 200 |
batchtopk |
k = 50 / 100 / 200 |
matryoshka_batchtopk |
k = 50 / 100 / 200 |
combinations.json lists all 162 training configurations. Two TinyStories configurations did not complete training and have no checkpoint: matryoshka_batchtopk_L0_65k_k100 and matryoshka_batchtopk_L0_65k_k200.
File layout
<base model>/checkpoints/<architecture>_L<layer>_<width>_<sparsity setting>/
sae_weights.safetensors # SAE parameters
sparsity.safetensors # feature sparsity statistics saved by SAELens
cfg.json # SAE configuration
runner_cfg.json # training configuration
Example: gpt2/checkpoints/topk_L6_16k_k100/.
Loading
The checkpoints are in SAELens on-disk format:
from huggingface_hub import snapshot_download
from sae_lens import SAE
root = snapshot_download(
"ritengzhang/SAE-AIU",
allow_patterns=["gpt2/checkpoints/topk_L6_16k_k100/*"],
)
sae = SAE.load_from_disk(f"{root}/gpt2/checkpoints/topk_L6_16k_k100", device="cpu")
The code repository's benchmark_training/load_benchmark_sae.py loads a checkpoint together with its base model.
Training data
GPT-2 and OPT-125M SAEs were trained on apollo-research/Skylion007-openwebtext-tokenizer-gpt2; TinyStories SAEs on apollo-research/roneneldan-TinyStories-tokenizer-gpt2. Each run uses 18,000 training steps at 4,096 tokens per batch. Full hyperparameters are in each runner_cfg.json.
License
The SAE checkpoints in this repository are released under the MIT License. The base models are not redistributed here and remain under their own licenses (GPT-2: MIT; OPT-125M: OPT model license, non-commercial research use; TinyStories model: no license declared on its model card).
Citation
@inproceedings{zhang2026rethinking,
title = {Rethinking SAE Evaluation: An Atomic Interpretable Unit Framework Reveals Hidden Polysemanticity and Redundancy},
author = {Zhang, Riteng and Wang, Xiaoqian},
booktitle = {Advances in Neural Information Processing Systems (Evaluations and Datasets Track)},
year = {2026}
}