LDM, 1.5B
LDM at 1.5B parameters from the paper Depth Memory, trained on 100B FineWeb-Edu tokens (step 94,972). The repository holds the final checkpoint of every run of the paper's learning-rate sweep for this model, one folder per peak learning rate; the paper reports the selected one. Code: https://github.com/davidgonmar/depth-memory.
| Folder | Peak learning rate | Held-out FineWeb-Edu cross-entropy |
|---|---|---|
lr5e-4/ |
5e-4 | 2.2207 |
lr1e-3/ |
1e-3 | 2.2141 (selected) |
lr2e-3/ |
2e-3 | 2.2260 |
Each folder is an OLMo-core checkpoint (config.json and
model_and_optim/) in the format of the code release, https://github.com/davidgonmar/depth-memory. Weights are BF16, the precision the training
forward ran in, except the parameters that training kept in FP32 (the mHC controllers and the LDM forgetting
parameters), which are FP32.
Usage
Install the code release, then download one folder:
huggingface-cli download davidgonmar/depth-memory-1.5b-ldm --include "lr1e-3/*" --local-dir 1.5b-ldm
Evaluate it as in the paper:
bash paper-experiments/pretrain/evaluation/run.sh 1.5b-ldm/lr1e-3 results.json
or load it in Python:
from olmo_core.generate import GenerationConfig, TransformerGenerationModule
module = TransformerGenerationModule.from_checkpoint(
"1.5b-ldm/lr1e-3", generation_config=GenerationConfig(pad_token_id=1, eos_token_id=50279)
)