math-grpo-cliphigh-step1000

GRPO + Clip-Higher baseline on ScienceQA, trained with GRPO on Qwen/Qwen2.5-Math-1.5B.

Selected as best validation pass@6 for this arm (rank 1).

Decoupled PPO clip bounds on MATH: lower 1-0.2, upper 1+0.28 (DAPO). One of DAPO's four components; not full DAPO.

Do not apply a chat template

Trained on raw prompt text. verl's RLHFDataset has apply_chat_template=False and it was never enabled. Applying Qwen2.5-Math's chat template at inference creates a train/eval mismatch measured at roughly 19 points of pass@1 on a sibling task.

from vllm import LLM, SamplingParams
llm = LLM(model="sandeep123/math-grpo-cliphigh-step1000", dtype="bfloat16", max_model_len=1536)
params = SamplingParams(n=6, temperature=1.0, top_p=1.0, top_k=-1, max_tokens=1024)
out = llm.generate([prompt_text], sampling_params=params)   # raw string, not llm.chat()

Validation metrics at this checkpoint

metric value
pass@1 0.6530
pass@6 0.9062
step 1000

Answer extraction (pre-registered). An answer is the content of the final \boxed{}; if absent, the last standalone A-E token. Responses with no extractable answer are scored incorrect, and all K rollouts stay in the denominator. This is ScienceQA's answer-choice accuracy, reported as "sampled answer accuracy (pass@1)".

Validation uses 256 held-out prompts, K=6, temperature 1.0, seed 42 -- pinned in code so every arm, including the temperature-1.2 arm, is scored under identical decoding.

Settings (identical across all baseline arms)

dataset ScienceQA (scienceqa_boxfix)
epochs / steps 25 / 1250
batch / rollouts 128 prompts, K=6
learning rate 1e-6 constant
KL (in-reward) 0.01
max prompt / response 512 / 1024 tokens
format reward 0.03, constant, no decay
seed 42

Arm-specific: entropy_coeff=0.0, clip=0.2/0.28, rollout temperature=1.0.

Note on checkpoint selection

Quality-optimal and diversity-optimal checkpoints differ substantially: for these runs the best-pass@1 checkpoint lands near step 1000-1200 while the best-pass@6 checkpoint is near step 200-500. Both are published so that diversity results are not reported from a checkpoint chosen purely for accuracy.

Downloads last month
12
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sandeep123/math-grpo-cliphigh-step1000

Finetuned
(237)
this model