openai/gsm8k
Benchmark β’ Updated β’ 17.6k β’ 944k β’ 1.51k
This model is fine-tuned using Group Relative Policy Optimization (GRPO) on 500 reasoning prompts from the GSM8K dataset as part of the AetherControl platform.
π GRPO Fine-Tuning Progression & Telemetry Log
βββββββββββ³ββββββββββββββ³ββββββββββββββ³ββββββββββββββ³ββββββββββββββ³βββββββββββββ
β Step β Mean Reward β Format β Accuracy β KL β GRPO Loss β
β β (r_mean) β Reward β Reward β Div (D_KL) β (L_grpo) β
β‘βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ©
β Step 01 β 0.42 β 0.23 β 0.22 β 0.0376 β 1.1934 β
β Step 05 β 0.50 β 0.34 β 0.32 β 0.0581 β 0.9610 β
β Step 10 β 0.64 β 0.54 β 0.48 β 0.0595 β 0.7322 β
β Step 15 β 0.72 β 0.75 β 0.70 β 0.0536 β 0.4367 β
β Step 20 β 0.85 β 0.92 β 0.87 β 0.0586 β 0.1730 β
βββββββββββ΄ββββββββββββββ΄ββββββββββββββ΄ββββββββββββββ΄ββββββββββββββ΄βββββββββββββ
<think> tags) and exact math string match (math_reward.py).Github Repository: https://github.com/Aravind0403/Building-and-Optimizing-Production-LLM-Serving-System