mlfoundations/dclm-baseline-1.0
Viewer • Updated • 3.02B • 545k • 319
Paper: Distributionally Robust Mixture-of-Experts Training
Project page: https://drmoet.github.io Code: https://github.com/MAPS-research/DRMoET
The weights in this repository are the checkpoints from the DRMoET paper: two DRMoET models (activation credit), each paired with the FLAME-MoE baseline trained with the same data, seed, and schedule. The folders below describe each checkpoint.
| Folder | Method | Experts (top-k) | Layers / hidden | Train iters | DRO (β, η, α, credit) |
|---|---|---|---|---|---|
DRMoET-1.7B |
DRMoET, activation credit | 64 (top-6) + shared | 18 / 2048 | 32,000 | 0.999, 0.1, 1.0, L2 activation norm |
FLAME-MoE-1.7B |
FLAME-MoE baseline | 64 (top-6) + shared | 18 / 2048 | 32,000 | — |
DRMoET-290M-32E |
DRMoET, activation credit | 32 (top-6) + shared | 9 / 1024 | 16,000 | 0.999, 0.001, 1.0, L2 activation norm |
FLAME-MoE-290M-32E |
FLAME-MoE baseline | 32 (top-6) + shared | 9 / 1024 | 16,000 | — |
All models use sequence length 2048, global batch size 1024, learning rate 3e-4, auxiliary load-balancing loss coefficient 0.01, and the EleutherAI/pythia-12b tokenizer, trained on DCLM. Each DRMoET checkpoint and its baseline share the same seed (1,234 for 1.7B and 3,407 for 290M-32E).