mHC, 400M

mHC at 400M parameters from the paper Depth Memory, trained on 50B FineWeb-Edu tokens (step 47,684). The repository holds the final checkpoint of every run of the paper's learning-rate sweep for this model, one folder per peak learning rate; the paper reports the selected one. Code: https://github.com/davidgonmar/depth-memory.

Folder Peak learning rate Held-out FineWeb-Edu cross-entropy
lr5e-4/ 5e-4 2.5291
lr1e-3/ 1e-3 2.5043
lr2e-3/ 2e-3 2.5041 (selected)

Each folder is an OLMo-core checkpoint (config.json and model_and_optim/) in the format of the code release, https://github.com/davidgonmar/depth-memory. Weights are BF16, the precision the training forward ran in, except the parameters that training kept in FP32 (the mHC controllers and the LDM forgetting parameters), which are FP32.

Usage

Install the code release, then download one folder:

huggingface-cli download davidgonmar/depth-memory-400m-mhc --include "lr2e-3/*" --local-dir 400m-mhc

Evaluate it as in the paper:

bash paper-experiments/pretrain/evaluation/run.sh 400m-mhc/lr2e-3 results.json

or load it in Python:

from olmo_core.generate import GenerationConfig, TransformerGenerationModule

module = TransformerGenerationModule.from_checkpoint(
    "400m-mhc/lr2e-3", generation_config=GenerationConfig(pad_token_id=1, eos_token_id=50279)
)
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train davidgonmar/depth-memory-400m-mhc

Collection including davidgonmar/depth-memory-400m-mhc