midtrain-bridge: Pythia-1B / 60B cosine C4 backbone
A C4-only 1B-parameter pretraining run to 60B tokens, published with all 17 intermediate
checkpoints along the trajectory. This is the W=0 reference arm of a study on when to
introduce a new data distribution (code) during pretraining; the checkpoints are the fork points
from which code-mixed branches are launched.
The run
| architecture | pythia-1b (GPT-NeoX), 1,011,781,632 parameters |
| tokenizer | Pythia / GPT-NeoX, padded vocab 50277 |
| data | C4 (en) only, no code, no math |
| tokens | 60B (30,518 steps) |
| sequence length | 2048 |
| global batch | 960 sequences = 1,966,080 tokens/step |
| LR schedule | cosine, peak 3e-4 โ min 3e-5 (floor is exactly 10% of peak) |
| warmup | 10% of the horizon = 6B tokens, linear |
| optimizer | AdamW |
| precision | bf16 autocast, fp32 master weights |
The LR is keyed to absolute token count, not step index, so a branch forked from any checkpoint continues the parent schedule with no re-warmup.
Checkpoints
trunk/ holds the end of warmup; branch_A/ continues 6B โ 60B.
| file | tokens | step | phase |
|---|---|---|---|
trunk/trunk_branchpoint.pt |
6.00B | 3052 | end of warmup, LR at peak |
trunk/trunk_6.00B_step3052.pt |
6.00B | 3052 | same point, snapshot name |
branch_A/branch_A_9.00B_step4578.pt |
9.00B | 4578 | cosine decay |
branch_A/branch_A_12.00B_step6104.pt |
12.00B | 6104 | |
branch_A/branch_A_15.00B_step7630.pt |
15.00B | 7630 | |
branch_A/branch_A_18.00B_step9156.pt |
18.00B | 9156 | |
branch_A/branch_A_21.00B_step10682.pt |
21.00B | 10682 | |
branch_A/branch_A_24.00B_step12208.pt |
24.00B | 12208 | |
branch_A/branch_A_27.00B_step13733.pt |
27.00B | 13733 | |
branch_A/branch_A_30.00B_step15259.pt |
30.00B | 15259 | |
branch_A/branch_A_33.00B_step16785.pt |
33.00B | 16785 | |
branch_A/branch_A_36.00B_step18311.pt |
36.00B | 18311 | |
branch_A/branch_A_42.00B_step21363.pt |
42.00B | 21363 | |
branch_A/branch_A_48.00B_step24415.pt |
48.00B | 24415 | |
branch_A/branch_A_54.00B_step27466.pt |
54.00B | 27466 | |
branch_A/branch_A_57.00B_step28992.pt |
57.00B | 28992 | |
branch_A/branch_A_58.50B_step29755.pt |
58.50B | 29755 | |
branch_A/branch_A_60.00B_step30518.pt |
60.00B | 30518 | final, LR at 3e-5 |
Every file carries full AdamW optimizer state (exp_avg, exp_avg_sq) except the 60B final,
which is weights-only. That is why the final is ~4 GB while the rest are ~12 GB: it is the
terminal checkpoint, meant for evaluation and fine-tuning rather than for continuing training.
The 4 GB file is complete and not truncated.
File format
Each .pt is a torch.save dict:
{
"model": state_dict, # litgpt GPT, GPT-NeoX layout
"optimizer": state_dict, # AdamW; ABSENT in the 60B final
"completed_steps": int,
"global_tokens": int, # absolute token count, keys the LR schedule
"config": dict, # full run config
"torch_rng": ..., "numpy_rng": ...,
"val_c4": float, "val_code": float, # held-out losses at that snapshot
}
Loading the weights:
import torch
ck = torch.load("branch_A/branch_A_60.00B_step30518.pt", map_location="cpu", weights_only=False)
print(ck["global_tokens"], ck["val_c4"])
state = ck["model"] # GPT-NeoX parameter layout
These are litgpt-format state dicts, not transformers checkpoints, so
AutoModelForCausalLM.from_pretrained will not read them directly. The parameter layout is
standard GPT-NeoX and converts mechanically.
Data
C4 (en), pre-tokenized, drawn from a 60.1B-token pool: the 40.1B pool published at
Impliedhomeland/midtrain-bridge-data
(pythia-70m/c4/, which serves the whole Pythia suite since all sizes share one tokenizer)
concatenated with 20.0B disjoint tokens from later C4 shards. Blocks are consumed in a fixed
seed-1 permutation, so the first 40.1B of this run's stream matches the published pool exactly.
Intended use
Released so the intro-timing experiments built on these fork points can be reproduced, and as a set of intermediate checkpoints along a single well-specified 1B run. This is a base model trained only on C4 with no instruction tuning, no safety filtering beyond C4's own, and no alignment work of any kind. Outputs will reflect whatever is in C4.