File size: 783 Bytes
32031f4 38e147f ae8b98f 3f6b5f0 3c56b4b 3f6b5f0 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 | ---
license: mit
---
[arXiv](https://arxiv.org/abs/2607.11505)
## 馃摝 Model Weights
Pre-trained GRPO checkpoints are available on [Huggingface](https://huggingface.co/KnowledgeXLab/PUST-Experiments):
| Checkpoint | Role | Training |
|:--|:--|:--|
| [`Qwen3-1.7B-Math-GRPO-Steps500`](https://huggingface.co/KnowledgeXLab/PUST-Experiments/tree/main/Qwen3-1.7B-Math-GRPO-Steps500) | Proxy | DeepMath-103K 路 GRPO 路 500 steps |
| [`Qwen3-1.7B-Math-GRPO-Steps800`](https://huggingface.co/KnowledgeXLab/PUST-Experiments/tree/main/Qwen3-1.7B-Math-GRPO-Steps800) | Proxy | DeepMath-103K 路 GRPO 路 800 steps |
| [`Qwen3-8B-Math-GRPO-Steps400`](https://huggingface.co/KnowledgeXLab/PUST-Experiments/tree/main/Qwen3-8B-Math-GRPO-Steps400) | Primary | DeepMath-103K 路 GRPO 路 400 steps | |