YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Phi-3.5-mini-instruct โ€” GT single-model GRPO on MATH345 (LLM side)

ground-truth answer reward. 136 steps = 1 epoch, 128 prompts/update, K=12, beta=0, lr 3e-6, bnpo loss, adam_beta2 0.95, eval every 10 steps. best/ = best-by-val (step 100); endpoint/ = step 136; training/ = train.log, trainer_state, best_metric.

Consolidated from the earlier split repos (-best / -endpoint). Local source: trl-projects-llm/work_dirs/stage/gt__microsoft_Phi-3p5-mini-instruct__20260804_230324

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support