slm-125m-ppo
PPO/RLAIF of the 125M legal/financial SFT model against a reward model (125M SFT + scalar head, Bradley-Terry on the preference triplets). 60-step PPO loop; mean reward -1.27 -> -0.64.
Base: rahulreddyhanu/slm-125m-legal-financial-sft. Part of an RLAIF demo (DPO + PPO) on small legal/financial models.
- Downloads last month
- 20
Model tree for rahulreddyhanu/slm-125m-ppo
Base model
thesreedath/slm-125m-base