YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Gemma-3-4B-it GT-GRPO on MATH345 (LLM side) โ two recipes, one repo
Both runs: 136 steps = 1 epoch, 128 prompts/update, K=12, beta=0, lr 3e-6, ground-truth answer reward. Gemma was NOT used in the final 3-agent lineup; kept for rebuttal (loss-type sensitivity, Gemma trainability).
- best/, endpoint/ (top level) = run2 canonical (bnpo, adam_beta2 0.95, eval every 10 steps). best-by-val is step 10 (eval reward 0.748); eval declines from 0.76 at step 0 to 0.726 at step 130.
- run2-bnpo-b2-0.95/checkpoint-{10..136}/ + best/: all 14 weight-only checkpoints of the canonical run (same recipe as all other llm-math345-* single-model runs).
- run1-dapo-b2-0.999/checkpoint-{10..136}/: 14 weight-only checkpoints of the earlier pilot run with dapo loss and adam_beta2 0.999, no eval during training. Different loss -> different experiment, not a duplicate.
- training/run1/, training/run2/: train.log, trainer_state, best_metric. Optimizer states are not uploaded (run1 step-136 optimizer retained locally). Local sources: trl-projects-llm/work_dirs/stage/gt__google_gemma-3-4b-it__20260803_190217 (run1), ..._20260804_111030 (run2). WandB project: grpo-training.
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support