MathaMAGICal checkpoints

Research checkpoints from MathaMAGICal (Algoverse, HCAI Lab GT / EleutherAI collaborators), the data-attribution study of how small GPT-2-style models trained from scratch acquire a held-out arithmetic composition (ADD then MULTIPLY). Every checkpoint is an HF-format GPT2LMHeadModel with the project's 276-token byte/digit tokenizer (tokenizer.json in each folder); context 512.

Load one with AutoModelForCausalLM.from_pretrained("leejason2025/mathamagical-checkpoints", subfolder="<folder>").

folder what it is
86m-cap2-300m-clip-s0/step_2375 86M, 300M-token run, before the skill (window winB2-2375 init)
86m-cap2-300m-clip-s0/step_2755 86M, 300M-token run, at the transition, 1/200 held-out (init of the headline windows)
86m-cap2-300m-clip-s0/step_3135 86M, 300M-token run, skill half-formed (winB2-3135 init)
86m-cap2-300m-clip-s0/step_5700 86M, 300M-token run, 144/200 held-out (C2-post5700 init)
86m-cap2-300m-clip-s0/step_9155 86M, 300M-token run, final, 177/200 held-out
86m-cap2-s0/step_2784 86M, 1B-token run, at its transition (50/200)
86m-cap2-s0/step_30517 86M, 1B-token run, final, 189/200 held-out
86m-cap2-muon-s0/step_870 86M, 1B-token Muon run, at its transition (18/200)
86m-cap2-muon-s0/step_30517 86M, 1B-token Muon run, final, 194/200 held-out
86m-nl86m-300m/step_9155 86M pretrained on 300M web-only tokens
86m-code86m-300m/step_9155 86M pretrained on 300M Python-only tokens
86m-CUR-nl-mathphase/step_700 web-first arm, math phase step 700 (pre-transition; CUR-nl-win init)
86m-CUR-code-mathphase/step_400 code-first arm, math phase step 400 (pre-transition; CUR-code-win init)
150m-web150m-1b/step_15312 150M web-only base at 50% of 1B tokens (branch point 15,312)
150m-web150m-1b/step_30517 150M web-only base, final (1B tokens)
150m-P150s-b15312-m1.0/step_400 150M, branch 15,312, 100% math phase step 400 (5/200; W150s-b15312-m1.0 init)
150m-P150s-b15312-m0.5/step_700 150M, branch 15,312, 50% math phase step 700 (2/200; W150s-b15312-m0.5 init)
150m-P150s-b30517-m1.0/step_400 150M, branch 30,517, 100% math phase step 400 (1/200; W150s-b30517-m1.0 init)
150m-P150s-b15312-m1.0/step_800 150M, branch 15,312, 100% math phase final (129/200 held-out)
150m-P150s-b30517-m1.0/step_800 150M, branch 30,517, 100% math phase final (141/200 held-out)
150m-P150s-b15312-m0.5/step_1600 150M, branch 15,312, 50% math phase final at step 1,600 (136/200 held-out)
86m-CUR-nl-mathphase/step_800 web-first arm, math phase final (26/200 held-out)
86m-CUR-code-mathphase/step_800 code-first arm, math phase final (130/200 held-out)
150m-mix150m-3b/step_1510 150M, 3B-token 50/50 run, before the skill, 0/200 (W3B-before init)
150m-mix150m-3b/step_2114 150M, 3B-token 50/50 run, skill appearing, 17/200 (W3B-onset init)
150m-mix150m-3b/step_2718 150M, 3B-token 50/50 run, skill forming, 89/200 (W3B-midrise init)
150m-mix150m-3b/step_4530 150M, 3B-token 50/50 run, skill present, 174/200 (W3B-after init)
150m-mix150m-3b/step_91552 150M, 3B-token 50/50 run, FINAL: 200/200 held-out, 100/100 in-distribution

Companion code and results: https://github.com/HCAI-Lab-GT/mathamagical (branch curriculum). Training data: leejason2025/mathamagical-3b-data (the 150M models) and leejason2025/magic-cap1b-data (the 86M models). Held-out accuracy numbers are exact match on 200 free-running ADD→MULTIPLY problems.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support