slm-125m-instruct-ppo

RLAIF (PPO (RLHF) with an AI-feedback reward model + critic) refinement of sudhisrk1982/slm-125m-instruct, trained on 500 AI-generated preference triplets (chosen from Gemini 2.5 Flash-Lite, rejected sampled from the SFT model itself, judged by Gemini 2.5 Flash).

Result: comparable to SFT, stable. Strict actor-critic PPO (value head + GAE, 60 rollout steps). Reward rose steadily and KL stayed leashed; output remains fluent, occasionally off-topic.

Not legal advice.

Downloads last month
252
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sudhisrk1982/slm-125m-instruct-ppo

Finetuned
(3)
this model