NBPO HelpSteer2 four-objective panel: MaxMin weighting

Stage 4 of a four-stage run on HelpSteer2 with four objectives (helpfulness, correctness, coherence, conciseness). Preferences come from a prompted pairwise oracle (Qwen3-32B, one attribute rubric per comparison, swap-averaged over presentation order); no reward model is fitted. Base and reference are meta-llama/Llama-3.1-8B-Instruct.

All rules in this panel share prompts, samples and judged comparisons, and differ only in the per-objective weight applied inside the same update. Every rule was carried to the same four stages.

Win rate against the reference on 500 held-out prompts, averaged over two evaluation judges (Phi-4 and Llama-3.3-70B), neither of which supplied the training preferences: worst objective 0.369, average 0.410. Single seed.

Downloads last month
12
Safetensors
Model size
8B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for promotion/nbpo-helpsteer2-maxmin-stage4

Finetuned
(2938)
this model