sqa-grpo-vanilla-step900
GRPO baseline on ScienceQA, trained with GRPO on Qwen/Qwen2.5-Math-1.5B.
Selected as best validation pass@1 for this arm (rank 1).
Plain GRPO. entropy_coeff=0 (the framework default of 1e-3 is explicitly disabled so "vanilla" means vanilla), symmetric PPO clip 0.2, rollout temperature 1.0.
Do not apply a chat template
Trained on raw prompt text. verl's RLHFDataset has apply_chat_template=False
and it was never enabled. Applying Qwen2.5-Math's chat template at inference creates
a train/eval mismatch measured at roughly 19 points of pass@1 on a sibling task.
from vllm import LLM, SamplingParams
llm = LLM(model="sandeep123/sqa-grpo-vanilla-step900", dtype="bfloat16", max_model_len=1536)
params = SamplingParams(n=6, temperature=1.0, top_p=1.0, top_k=-1, max_tokens=1024)
out = llm.generate([prompt_text], sampling_params=params) # raw string, not llm.chat()
Validation metrics at this checkpoint
| metric | value |
|---|---|
| pass@1 | 0.8568 |
| pass@6 | 0.9609 |
| step | 900 |
Answer extraction (pre-registered). An answer is the content of the final
\boxed{}; if absent, the last standalone A-E token. Responses with no
extractable answer are scored incorrect, and all K rollouts stay in the
denominator. This is ScienceQA's answer-choice accuracy, reported as
"sampled answer accuracy (pass@1)".
Validation uses 256 held-out prompts, K=6, temperature 1.0, seed 42 -- pinned in code so every arm, including the temperature-1.2 arm, is scored under identical decoding.
Settings (identical across all baseline arms)
| dataset | ScienceQA (scienceqa_boxfix) |
| epochs / steps | 25 / 1250 |
| batch / rollouts | 128 prompts, K=6 |
| learning rate | 1e-6 constant |
| KL (in-reward) | 0.01 |
| max prompt / response | 512 / 1024 tokens |
| format reward | 0.03, constant, no decay |
| seed | 42 |
Arm-specific: entropy_coeff=0.0, clip=0.2/0.2, rollout temperature=1.0.
Note on checkpoint selection
Quality-optimal and diversity-optimal checkpoints differ substantially: for these runs the best-pass@1 checkpoint lands near step 1000-1200 while the best-pass@6 checkpoint is near step 200-500. Both are published so that diversity results are not reported from a checkpoint chosen purely for accuracy.
- Downloads last month
- -