MathaMAGICal checkpoints
Research checkpoints from MathaMAGICal (Algoverse, HCAI Lab GT / EleutherAI collaborators), the data-attribution
study of how small GPT-2-style models trained from scratch acquire a held-out arithmetic composition
(ADD then MULTIPLY). Every checkpoint is an HF-format GPT2LMHeadModel with the project's 276-token
byte/digit tokenizer (tokenizer.json in each folder); context 512.
Load one with AutoModelForCausalLM.from_pretrained("leejason2025/mathamagical-checkpoints", subfolder="<folder>").
| folder | what it is |
|---|---|
86m-cap2-300m-clip-s0/step_2375 |
86M, 300M-token run, before the skill (window winB2-2375 init) |
86m-cap2-300m-clip-s0/step_2755 |
86M, 300M-token run, at the transition, 1/200 held-out (init of the headline windows) |
86m-cap2-300m-clip-s0/step_3135 |
86M, 300M-token run, skill half-formed (winB2-3135 init) |
86m-cap2-300m-clip-s0/step_5700 |
86M, 300M-token run, 144/200 held-out (C2-post5700 init) |
86m-cap2-300m-clip-s0/step_9155 |
86M, 300M-token run, final, 177/200 held-out |
86m-cap2-s0/step_2784 |
86M, 1B-token run, at its transition (50/200) |
86m-cap2-s0/step_30517 |
86M, 1B-token run, final, 189/200 held-out |
86m-cap2-muon-s0/step_870 |
86M, 1B-token Muon run, at its transition (18/200) |
86m-cap2-muon-s0/step_30517 |
86M, 1B-token Muon run, final, 194/200 held-out |
86m-nl86m-300m/step_9155 |
86M pretrained on 300M web-only tokens |
86m-code86m-300m/step_9155 |
86M pretrained on 300M Python-only tokens |
86m-CUR-nl-mathphase/step_700 |
web-first arm, math phase step 700 (pre-transition; CUR-nl-win init) |
86m-CUR-code-mathphase/step_400 |
code-first arm, math phase step 400 (pre-transition; CUR-code-win init) |
150m-web150m-1b/step_15312 |
150M web-only base at 50% of 1B tokens (branch point 15,312) |
150m-web150m-1b/step_30517 |
150M web-only base, final (1B tokens) |
150m-P150s-b15312-m1.0/step_400 |
150M, branch 15,312, 100% math phase step 400 (5/200; W150s-b15312-m1.0 init) |
150m-P150s-b15312-m0.5/step_700 |
150M, branch 15,312, 50% math phase step 700 (2/200; W150s-b15312-m0.5 init) |
150m-P150s-b30517-m1.0/step_400 |
150M, branch 30,517, 100% math phase step 400 (1/200; W150s-b30517-m1.0 init) |
150m-P150s-b15312-m1.0/step_800 |
150M, branch 15,312, 100% math phase final (129/200 held-out) |
150m-P150s-b30517-m1.0/step_800 |
150M, branch 30,517, 100% math phase final (141/200 held-out) |
150m-P150s-b15312-m0.5/step_1600 |
150M, branch 15,312, 50% math phase final at step 1,600 (136/200 held-out) |
86m-CUR-nl-mathphase/step_800 |
web-first arm, math phase final (26/200 held-out) |
86m-CUR-code-mathphase/step_800 |
code-first arm, math phase final (130/200 held-out) |
150m-mix150m-3b/step_1510 |
150M, 3B-token 50/50 run, before the skill, 0/200 (W3B-before init) |
150m-mix150m-3b/step_2114 |
150M, 3B-token 50/50 run, skill appearing, 17/200 (W3B-onset init) |
150m-mix150m-3b/step_2718 |
150M, 3B-token 50/50 run, skill forming, 89/200 (W3B-midrise init) |
150m-mix150m-3b/step_4530 |
150M, 3B-token 50/50 run, skill present, 174/200 (W3B-after init) |
150m-mix150m-3b/step_91552 |
150M, 3B-token 50/50 run, FINAL: 200/200 held-out, 100/100 in-distribution |
Companion code and results: https://github.com/HCAI-Lab-GT/mathamagical (branch curriculum).
Training data: leejason2025/mathamagical-3b-data (the 150M models) and
leejason2025/magic-cap1b-data (the 86M models).
Held-out accuracy numbers are exact match on 200 free-running ADD→MULTIPLY problems.