YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

To Mix or To Merge?
Toward Multi-Domain Reinforcement Learning for Large Language Models

arXiv Hugging Face ModelScope GitHub COLM 2026

Haoqing Wang†, Xiang Long†, Ziheng Li†, Yilong Xu, Tingguang Li, Yehui Tangβœ‰
Samsung Research, Beijing, China   Β·   Peking University


πŸ“° News

  • [2026.09.07] πŸŽ‰ The model checkpoints are now open-sourced on Hugging Face and ModelScope! Feel free to use our checkpoints for your post-training research (e.g., weight merging, multi-teacher on-policy distillation)!
  • [2026.07.09] πŸŽ‰ Our paper is accepted to COLM 2026!

πŸ“š Citation

If you find this work useful, please consider citing:

@inproceedings{
wang2026to,
title={To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models},
author={Haoqing Wang and Xiang Long and Ziheng Li and Yilong Xu and Tingguang Li and Yehui Tang},
booktitle={Third Conference on Language Modeling},
year={2026},
url={https://openreview.net/forum?id=jP7j5XkG8J}
}
Downloads last month
11
Safetensors
Model size
4B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Collection including Jackwang111/M2RL-MT_OPD

Paper for Jackwang111/M2RL-MT_OPD