Safetensors
File size: 783 Bytes
32031f4
 
 
38e147f
ae8b98f
 
3f6b5f0
 
3c56b4b
3f6b5f0
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
---
license: mit
---

[arXiv](https://arxiv.org/abs/2607.11505)

## 馃摝 Model Weights

Pre-trained GRPO checkpoints are available on [Huggingface](https://huggingface.co/KnowledgeXLab/PUST-Experiments):

| Checkpoint | Role | Training |
|:--|:--|:--|
| [`Qwen3-1.7B-Math-GRPO-Steps500`](https://huggingface.co/KnowledgeXLab/PUST-Experiments/tree/main/Qwen3-1.7B-Math-GRPO-Steps500) | Proxy | DeepMath-103K 路 GRPO 路 500 steps |
| [`Qwen3-1.7B-Math-GRPO-Steps800`](https://huggingface.co/KnowledgeXLab/PUST-Experiments/tree/main/Qwen3-1.7B-Math-GRPO-Steps800) | Proxy | DeepMath-103K 路 GRPO 路 800 steps |
| [`Qwen3-8B-Math-GRPO-Steps400`](https://huggingface.co/KnowledgeXLab/PUST-Experiments/tree/main/Qwen3-8B-Math-GRPO-Steps400) | Primary | DeepMath-103K 路 GRPO 路 400 steps |