code-q25_3b-nobuf

GRPO on MBPP from Qwen/Qwen2.5-3B. Arm: GRPO, no replay buffer.

Part of a study of forgetting during code RL: the same run is trained with and without an SFT-replay buffer, and evaluated on MBPP+ across training.

Layout

One subfolder per checkpoint, global_step_<N>/, each a full Hugging Face model directory. Steps are the stride-30 evaluation grid plus the run's endpoint (30, 60, 90, 120, 150, 180, 210, 240, 270, 300).

from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("RL-Forgetting-Experiments-3/code-q25_3b-nobuf", subfolder="global_step_300")
t = AutoTokenizer.from_pretrained("RL-Forgetting-Experiments-3/code-q25_3b-nobuf", subfolder="global_step_300")

Training

  • Data: 320 MBPP problems, held out from both MBPP+ (378) and MBPP's canonical test split (276), so both remain reportable.
  • GRPO, 8 rollouts per prompt, batch 64, actor lr 1e-6, response length 3072.
  • Reward: execution of the generated program against the task's asserts (binary).
  • Buffer arms replay 128 past rollouts per step with weight lambda=0.1, sampled by hard_cooldown.

Evaluation

MBPP+ (378 problems), n=160 samples, temperature 0.6, top_p 0.95, unbiased pass@k. Note that scoring against MBPP+'s full test suite is substantially stricter than MBPP's original 3 asserts (~8-10 points of pass@1).

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for RL-Forgetting-Experiments-3/code-q25_3b-nobuf

Base model

Qwen/Qwen2.5-3B
Finetuned
(596)
this model

Collection including RL-Forgetting-Experiments-3/code-q25_3b-nobuf