paper-filter-metric-1.5b

Recipe: recipes/papers/filter-metric · Collection: Papers, replicated

GRPO on GSM8K with a length-shaped reward, where the all-equal group filter reads the binary outcome instead of the shaped score. Filtering on the shaped score keeps all-wrong groups alive and turns length crumbs into full-size advantages. Filtering on the outcome drops them, which is what DAPO's dynamic sampling means.

Result

Arm pass@1 95% CI pass@k Steps GPU min
Base, no training 0.36 [0.29, 0.43] 0.57 0 0
Baseline (filter on shaped score) 0.39 [0.33, 0.45] 0.68 40 30.0
Recipe (filter on binary outcome) 0.46 [0.39, 0.53] 0.70 40 21.8

Recipe vs baseline: +0.067 [+0.021, +0.113] over 120 paired tasks. The interval excludes zero and the delta clears the eval's re-run band (0.025 from ten base re-runs). The shaped score went down while pass@1 went up, the opposite of over-optimization. Verdict: unresolved until a second training seed per arm.

Arms in this repo

The root holds the arm the recipe README's headline number reports. Every other arm is a subfolder named after it. checkpoints/ never ships.

folder arm
. recipe arm: filter on binary outcome, 2026-09-17 run
baseline baseline arm: filter on shaped score, 2026-09-17 run

Load

from peft import PeftModel
from transformers import AutoModelForCausalLM

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
model = PeftModel.from_pretrained(base, "while-ai/paper-filter-metric-1.5b")  # the headline arm
model = PeftModel.from_pretrained(base, "while-ai/paper-filter-metric-1.5b", subfolder="baseline")  # another arm

Reproduce

git clone https://github.com/whilehq/whileai-sdk && cd whileai-sdk/recipes/papers/filter-metric
python recipe.py

The recipe README pins the seed, the library versions and the GPU, and its Checks table says what the eval verified. Read the Learned section before quoting a number from this card.

Downloads last month
26
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for while-ai/paper-filter-metric-1.5b

Adapter
(1483)
this model

Dataset used to train while-ai/paper-filter-metric-1.5b

Collection including while-ai/paper-filter-metric-1.5b