TIPS-Qwen3-4B-Instruct-2507-Math

This TIPS math checkpoint is initialized from Qwen/Qwen3-4B-Instruct-2507 and trained with outcome-only GRPO. It is a generative reward model that reasons over a mathematical solution before producing step-level and outcome labels.

Training data and code are available at https://huggingface.co/datasets/XingYing-stack/TIPS-Training-Data and https://github.com/RUCBM/TIPS.

Use the prompt templates and evaluation scripts in the TIPS repository. This checkpoint is intended for reward modeling and process verification rather than general-purpose chat.

Built upon verl and released under Apache-2.0.

Downloads last month
-
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for XingYing-stack/TIPS-Qwen3-4B-Instruct-2507-Math

Finetuned
(1957)
this model

Dataset used to train XingYing-stack/TIPS-Qwen3-4B-Instruct-2507-Math