slm-125m-instruct-ppo
RLAIF (PPO (RLHF) with an AI-feedback reward model + critic) refinement of sudhisrk1982/slm-125m-instruct,
trained on 500 AI-generated preference triplets (chosen from
Gemini 2.5 Flash-Lite, rejected sampled from the SFT model itself, judged by
Gemini 2.5 Flash).
Result: comparable to SFT, stable. Strict actor-critic PPO (value head + GAE, 60 rollout steps). Reward rose steadily and KL stayed leashed; output remains fluent, occasionally off-topic.
Not legal advice.
- Downloads last month
- 252
Model tree for sudhisrk1982/slm-125m-instruct-ppo
Base model
thesreedath/slm-125m-base Finetuned
sudhisrk1982/slm-125m-instruct