cs2881-hw1-combined
Combined RLAIF+RLVR — Setting B continued with GRPO on a within-batch 50/50 mix of the persona judge and the OpenBookQA verifier (per-type normalized reward). Both axes move (OpenBookQA +0.024, clf_style +0.257, register +1.105). Merged model at the repo root; the LoRA adapter is in adapter/.
- Downloads last month
- -
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support