RLAIF — Setting B continued with GRPO using a gpt-4o-mini Sheldon-adherence judge as the reward. Persona up (clf_style 0.571 -> 0.876, register 1.569 -> 2.108), capability held. Merged model at the repo root; the LoRA adapter is in adapter/.
adapter/
Chat template
Files info
Base model