A Controlled Study of Attention-Only Transformers
Paper • 2607.18363 • Published • 5
Simple Attention Network — a ~31M-parameter custom LLM, a faithful PyTorch port of the
needle architecture (arXiv:2607.18363). Trained locally on an RTX 4070 Ti 12 GB.
This repo contains the released weights for the 12-layer configuration.
| file | what |
|---|---|
pytorch_model.bin |
model state_dict (load via san_model.SimpleAttentionNetwork.load_state_dict) |
san_latest.pt |
raw training checkpoint (step 137,260): {step, loss, model_state_dict, optimizer_state_dict, config} |
config.json |
architecture hyperparameters |
tokenizer/ |
SmolLM2-135M tokenizer (vocab 49152) |
san_model.py, san_triton.py |
model definition (needed to load the weights) |
load_example.py |
minimal load + forward example |
from san_model import SimpleAttentionNetwork, SANConfig
import torch
cfg = SANConfig(num_layers=12)
model = SimpleAttentionNetwork(cfg).eval()
sd = torch.load("pytorch_model.bin", map_location="cpu")
model.load_state_dict(sd)
See load_example.py. Full training/eval code: GitHub
kenpeter/x-small.
benchmarks/ppl_trend.png, data in benchmarks/ppl_trend.csv): code/math/prose perplexity across the 12-layer evals (steps 63,663 → 110,938 → 123,141).benchmarks/throughput.png, data in benchmarks/tput_release.csv): released 12-layer weights, eager forward+backward with grad-checkpointing, bf16, seq 2048, RTX 4070 Ti 12GB. B=8 OOM because the GPU was shared with another training job at bench time; the release run used B=8 on a dedicated GPU.MIT (code). Weights released for research use.