Math-Qwen3-1.7B-Qwen3-4B-Non-Thinking-RL-Math-Step500

Paper · GitHub · Project · Collection

This repository contains an on-policy distillation (OPD) student checkpoint released with Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling. Its base model, teacher, and training data are listed below.

Overview: pass@K gains depend on the sampling budget

Math pass@K results across settings

The figure reports paper results across math settings; it is not a scorecard for this checkpoint alone.

Checkpoint details

Step500 is part of the teacher's name, not a stated student training step.

More information

See the paper and GitHub repository for training settings, evaluation protocols, and analysis. Upstream model and data terms remain applicable; see third-party notices.

Citation

@misc{ge2026understandingonpolicydistillationlens,
  title={Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling},
  author={Xinmu Ge and Zizhuo Zhang and Yu Huang and Jianing Zhu and Lin Yuan and Wanli Gu and Weichang Wu and Weiran Huang and Bo Han and Xiaolu Zhang and Jiangchao Yao},
  year={2026},
  eprint={2608.11829},
  archivePrefix={arXiv},
  primaryClass={cs.LG},
  doi={10.48550/arXiv.2608.11829},
  url={https://arxiv.org/abs/2608.11829}
}
Downloads last month
215
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Geraldxm/Math-Qwen3-1.7B-Qwen3-4B-Non-Thinking-RL-Math-Step500

Finetuned
Qwen/Qwen3-1.7B
Finetuned
(1272)
this model

Collection including Geraldxm/Math-Qwen3-1.7B-Qwen3-4B-Non-Thinking-RL-Math-Step500

Paper for Geraldxm/Math-Qwen3-1.7B-Qwen3-4B-Non-Thinking-RL-Math-Step500