OLMo-20M checkpoint pretrained with SAM optimizer (rho 5e-2) and cosine LR schedule on 64B tokens from DCLM.

Paper:"Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting" Ishaan Watts*, Catherine Li*, Sachin Goyal, Jacob Mitchell Springer, Aditi Raghunathan ICML2026 https://arxiv.org/abs/2605.02105

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including WattsIshaan/OLMo-20m-64B-sam-cosine

Paper for WattsIshaan/OLMo-20m-64B-sam-cosine