LDM, D_k = 4, 400M

LDM, D_k = 4 at 400M parameters from the paper Depth Memory, trained on 50B FineWeb-Edu tokens (step 47,684). The repository holds the final checkpoint of every run of the paper's learning-rate sweep for this model, one folder per peak learning rate; the paper reports the selected one. Code: https://github.com/davidgonmar/depth-memory.

Folder Peak learning rate Held-out FineWeb-Edu cross-entropy
lr5e-4/ 5e-4 2.5178
lr1e-3/ 1e-3 2.4965 (selected)
lr2e-3/ 2e-3 2.4991

Each folder is an OLMo-core checkpoint (config.json and model_and_optim/) in the format of the code release, https://github.com/davidgonmar/depth-memory. Weights are BF16, the precision the training forward ran in, except the parameters that training kept in FP32 (the mHC controllers and the LDM forgetting parameters), which are FP32.

Usage

Install the code release, then download one folder:

huggingface-cli download davidgonmar/depth-memory-400m-ldm-dk4 --include "lr1e-3/*" --local-dir 400m-ldm-dk4

Evaluate it as in the paper:

bash paper-experiments/pretrain/evaluation/run.sh 400m-ldm-dk4/lr1e-3 results.json

or load it in Python:

from olmo_core.generate import GenerationConfig, TransformerGenerationModule

module = TransformerGenerationModule.from_checkpoint(
    "400m-ldm-dk4/lr1e-3", generation_config=GenerationConfig(pad_token_id=1, eos_token_id=50279)
)
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train davidgonmar/depth-memory-400m-ldm-dk4

Collection including davidgonmar/depth-memory-400m-ldm-dk4