sqa-grpo-cliphigh-step1000
GRPO + Clip-Higher baseline on ScienceQA, trained with GRPO on Qwen/Qwen2.5-Math-1.5B.
Selected as best validation pass@1 for this arm (rank 1).
Decoupled PPO clip bounds: lower 1-0.2, upper 1+0.28 (DAPO). Raising clip_ratio alone widens BOTH bounds and is a different intervention. Clip-Higher is one of DAPO's four components; this arm is not full DAPO.
Training data and format reward
| dataset | ScienceQA (scienceqa) |
| epochs / steps | 25 / 1250 |
| batch / rollouts | 128 prompts, K=6 |
| learning rate | 1e-6 constant |
| KL (in-reward) | 0.01 |
| max prompt / response | 512 / 1024 tokens |
| format reward | 0.03, constant, no decay |
| seed | 42 |
The training prompt requests the numbered-step scaffold and a final \boxed{X}, matching the format reward. It does not carry the "IMPORTANT rules" emphasis block used by the AoPS suite, so the wording (not the requested output format) differs from that suite.
Arm-specific: entropy_coeff=0.0, clip=0.2/0.28, rollout temperature=1.0.
Do not apply a chat template
Trained on raw prompt text. verl's RLHFDataset has apply_chat_template=False
and it was never enabled. Applying Qwen2.5-Math's chat template at inference creates
a train/eval mismatch measured at roughly 19 points of pass@1.
from vllm import LLM, SamplingParams
llm = LLM(model="sandeep123/sqa-grpo-cliphigh-step1000", dtype="bfloat16", max_model_len=1536)
params = SamplingParams(n=6, temperature=1.0, top_p=1.0, top_k=-1, max_tokens=1024)
out = llm.generate([prompt_text], sampling_params=params) # raw string, not llm.chat()
Validation metrics at this checkpoint
| metric | value |
|---|---|
| pass@1 | 0.8581 |
| pass@6 | 0.9766 |
| step | 1000 |
Answer extraction (pre-registered). An answer is the content of the final \boxed{}; if absent, the last standalone A-E token. Responses with no extractable answer are scored incorrect, and all K rollouts
stay in the denominator.
Validation uses 256 held-out prompts, K=6, temperature 1.0, seed 42 -- pinned in code so every arm, including the temperature-1.2 arm, is scored under identical decoding.
Checkpoint selection
Quality-optimal and diversity-optimal checkpoints differ substantially, so both the best-pass@1 and best-pass@6 checkpoints are published for every arm. Diversity results are therefore never reported from a checkpoint chosen purely for accuracy.
- Downloads last month
- 9