Text-pretrained baselines for Self-Play Pretraining with Zero Data
These are byte-level language models trained on ordinary web text (DCLM-Baseline). They use exactly the architecture and model sizes of the self-play learners from Self-Play Pretraining with Zero Data (weights). I trained them to answer a question the paper leaves open: does ordinary pretraining produce the same in-context learning that self-play does, at the same size and the same number of training tokens?
Short answer: no. Each gets a different kind. The full comparison, code and figures are at github.com/mihir-s-05/icl-selfplay-vs-text. Every ICL score is in the companion dataset rihim/icl-selfplay-vs-text-results.
What's in the repo
dclm/<size>/seed-<s>/ trained from scratch on DCLM text
sp2dclm/<size>/seed-0/ started from the paper's self-play checkpoint (round 8191), then DCLM
config.json model architecture ({"learner": {...}})
learner_<tokens>M.pth {"learner_state_dict", "tokens", "round", "arm", "init_from"}
log.jsonl training log: loss, LR, grad norm, validation bits/byte
run.json full training configuration
This is the same layout, file format and model class as the paper's own checkpoints, so anything that loads theirs loads these.
Checkpoint schedule.
- DCLM runs save at 0.1, 0.25, 0.5 and 1B tokens, then at 1.6, 3.2, 6.4, 12.9 and 17.7B. The last five are exactly the learner-token counts of the paper's self-play rounds 256, 512, 1024, 2048 and 2816 (1536 programs × 4096 bytes per round). That makes checkpoint n here directly comparable to self-play round n. Round 2816 (17.7B tokens) is the one the paper uses for its ICL results.
- Warm starts save at 0.05, 0.1, 0.25 and 0.5B tokens of DCLM.
The runs
| Size | Non-emb. params | Arch (d / heads / layers) | DCLM seeds | LR | Val bits/byte @ 17.7B | Warm start: val bits/byte @ 0.5B (scratch @ 0.5B) |
|---|---|---|---|---|---|---|
| 100k | 65,728 | 64 / 1 / 1 | 1 | 2e-3 | 2.495 | 2.786 (2.662) |
| 500k | 492,160 | 128 / 2 / 2 | 1 | 4e-3 | 1.882 | 2.116 (2.205) |
| 1M | 984,192 | 128 / 2 / 4 | 2 | 2e-3 | 1.711, 1.702 | 1.999 (2.099) |
| 3M | 3,016,960 | 256 / 4 / 4 | 2 | 2e-3 | 1.584, 1.581 | 1.821 (1.984) |
| 6M | 6,033,664 | 256 / 4 / 8 | 2 | 2e-3 | 1.506, 1.504 | 1.741 (1.778) |
| 24M | 24,253,184 | 512 / 8 / 8 | 1 | 5e-4 | 1.354 | 1.560 (1.666) |
Validation is on 64 sequences of a DCLM sample disjoint from training. The paper's own self-play learners score about 6.2 to 7.6 bits/byte on the same kind of text zero-shot. Warm starts beat scratch at equal tokens at every size except 100k, consistent with the paper's Fig. 6.
Training.
- Data: 18.5 GB of DCLM-Baseline 1.0 text, skipping every file the paper's evaluation
corpus draws from. Each sequence is the byte
Ofollowed by 4095 text bytes, which is the paper's scoring convention. - Optimiser and schedule:
- AdamW (β 0.9/0.95), weight decay 0.1 on all parameters.
- Batch 64 × 4096 bytes.
- 2% linear warmup, then constant LR (no decay), so every checkpoint is a plain mid-run model, like the self-play ones.
- bf16 autocast and
torch.compile, on one A100 80GB.
- LR choice: picked per size from a 3-point sweep of short runs (0.3 to 0.5B tokens).
- Initialisation: GPT-2 style, normal(0, 0.02) with scaled residual projections.
- Warm starts: LR 3e-3, the value the paper tuned for its own warm starts.
Loading one
You need the model class from the paper's repo (scoring/src/framework/model.py):
import json, torch
from huggingface_hub import hf_hub_download
from src.framework.model import ProgramLanguageModel # from nourya-aliz/self_play_pretraining/scoring
repo = "rihim/icl-selfplay-vs-text-checkpoints"
cfg = json.load(open(hf_hub_download(repo, "dclm/6M/seed-0/config.json")))["learner"]
blob = torch.load(hf_hub_download(repo, "dclm/6M/seed-0/learner_17723M.pth"),
map_location="cpu", weights_only=False)
model = ProgramLanguageModel(**cfg)
model.load_state_dict(blob["learner_state_dict"], strict=False) # same call as for the paper's checkpoints
model.eval()
text = b"O" + "The capital of France is".encode() # every sequence starts with the byte 'O'
ids = torch.tensor([list(text)])
next_byte = model(ids)[0][0, -1].argmax().item()
To get everything: hf download rihim/icl-selfplay-vs-text-checkpoints --local-dir ckpts.
Headline results (in-context learning at 17.7B tokens)
Mean exact-match accuracy over each suite's tasks. Self-play is the paper's checkpoints (4 seeds, 95% CI); DCLM is these models (one number per seed).
| Size | Printable-text tasks: self-play | Printable-text tasks: DCLM | Word tasks: self-play | Word tasks: DCLM |
|---|---|---|---|---|
| 1M | 0.30 ± 0.15 | 0.17, 0.13 | 0.31 ± 0.06 | 0.51, 0.52 |
| 3M | 0.27 ± 0.16 | 0.15, 0.15 | 0.36 ± 0.08 | 0.63, 0.51 |
| 6M | 0.44 ± 0.04 | 0.24, 0.14 | 0.40 ± 0.08 | 0.59, 0.60 |
| 24M | 0.41 ± 0.10 | 0.15 | 0.41 ± 0.06 | 0.69 |
What this shows:
- Self-play wins procedural tasks: copying by position, stack tracking, comparisons.
- These text models win lookup and meaning: key→value recall, and categorising words they haven't been shown.
- Both gaps widen with size.
Details, caveats and per-task numbers are in the write-up.
Limitations
- Small models: these are tiny research models, not useful as general language models.
- Seeds: there are only 1 to 2 seeds per size.
- Schedule: the LR is constant (no cooldown), so the final checkpoints are not fully converged for their token budget.
Credit
The architecture, model class and self-play method are from Self-Play Pretraining with Zero Data (arXiv 2609.30063). Training data is DCLM-Baseline 1.0.