Salesforce/LLaMA-3-8B-SFR-Iterative-DPO-R
Text Generation • 8B • Updated • 104 • • 78
None defined yet.
Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation