DRMoET checkpoints

Paper: Distributionally Robust Mixture-of-Experts Training

Project page: https://drmoet.github.io Code: https://github.com/MAPS-research/DRMoET

The weights in this repository are the checkpoints from the DRMoET paper: two DRMoET models (activation credit), each paired with the FLAME-MoE baseline trained with the same data, seed, and schedule. The folders below describe each checkpoint.

Checkpoints

Folder Method Experts (top-k) Layers / hidden Train iters DRO (β, η, α, credit)
DRMoET-1.7B DRMoET, activation credit 64 (top-6) + shared 18 / 2048 32,000 0.999, 0.1, 1.0, L2 activation norm
FLAME-MoE-1.7B FLAME-MoE baseline 64 (top-6) + shared 18 / 2048 32,000 —
DRMoET-290M-32E DRMoET, activation credit 32 (top-6) + shared 9 / 1024 16,000 0.999, 0.001, 1.0, L2 activation norm
FLAME-MoE-290M-32E FLAME-MoE baseline 32 (top-6) + shared 9 / 1024 16,000 —

All models use sequence length 2048, global batch size 1024, learning rate 3e-4, auxiliary load-balancing loss coefficient 0.01, and the EleutherAI/pythia-12b tokenizer, trained on DCLM. Each DRMoET checkpoint and its baseline share the same seed (1,234 for 1.7B and 3,407 for 290M-32E).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train TonyTeng/DRMoET

Paper for TonyTeng/DRMoET