Coding RL Checkpoints
Collection
6 items • Updated
GRPO on MBPP from Qwen/Qwen2.5-3B. Arm: GRPO, no replay buffer.
Part of a study of forgetting during code RL: the same run is trained with and without an SFT-replay buffer, and evaluated on MBPP+ across training.
One subfolder per checkpoint, global_step_<N>/, each a full Hugging Face model
directory. Steps are the stride-30 evaluation grid plus the run's endpoint
(30, 60, 90, 120, 150, 180, 210, 240, 270, 300).
from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("RL-Forgetting-Experiments-3/code-q25_3b-nobuf", subfolder="global_step_300")
t = AutoTokenizer.from_pretrained("RL-Forgetting-Experiments-3/code-q25_3b-nobuf", subfolder="global_step_300")
hard_cooldown.MBPP+ (378 problems), n=160 samples, temperature 0.6, top_p 0.95, unbiased pass@k. Note that scoring against MBPP+'s full test suite is substantially stricter than MBPP's original 3 asserts (~8-10 points of pass@1).
Base model
Qwen/Qwen2.5-3B