HOPE β€” Nested Learning in PyTorch

The index card for nested-learning-hope, an independent PyTorch reproduction of HOPE from "Nested Learning: The Illusion of Deep Learning Architectures" (arXiv:2512.24695, Behrouz et al., Google Research). No official code or weights exist for the paper; this is an unofficial reproduction at reduced scale.

HOPE is an attention-free language model whose weights change during the forward pass (test-time learning): Self-Modifying Titans (Eq. 83-93, gated delta rule, self-generated targets) plus a Continuum Memory System (Eq. 70-71, MLP levels updated every 1/2/4 chunks by the accumulated task-loss gradient).

The released recipe (reproducible in ~7 GPU-hours)

Parameters 30.6M (4 layers, d=512, chunk 512, no positional embedding)
Tokenizer none β€” raw bytes (vocab 256)
Training data FineWeb-Edu sample, 0.3B bytes, sequence length 2048
Held-out validation 1.721 bits/byte
Needle recall (1K / 6.6K bytes) +2.59 / +2.10 bpb vs control (n=40)
8K extrapolation (trained at 2K) flat, no horizon (1.66 β†’ 1.96 bpb)

Every number is reproducible from the GitHub repository (committed training log, evaluation scripts, and an exactness test verifying the chunk-parallel path equals the sequential recurrence to fp64 machine precision).

Reproducing the model

No weights are hosted here (yet) β€” the checkpoint reproduces from the GitHub repository in about 7 hours on one 16 GB GPU, and the committed training log lets you verify your run against ours at step 100/200 before committing the full budget:

git clone https://github.com/smallhours19/nested-learning-hope && cd nested-learning-hope
pip install -r requirements.txt
python scripts/download_data.py --out data/fineweb-edu --shards 3
python scripts/train_byte.py --data data/fineweb-edu --out runs/hope-byte

Where this implementation had to choose beyond the paper's text, every choice is documented in DEVIATIONS.md; stability devices the paper does not mention are cataloged in STABILITY.md.

Citation

@article{behrouz2025nested,
  title   = {Nested Learning: The Illusion of Deep Learning Architectures},
  author  = {Behrouz, Ali and Razaviyayn, Meisam and Zhong, Peilin and
             Mirrokni, Vahab},
  journal = {arXiv preprint arXiv:2512.24695},
  year    = {2025}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for smallhours19/nested-learning-hope