NBPO HelpSteer2 four-objective panel: MaxMin weighting
Stage 4 of a four-stage run on HelpSteer2 with four objectives (helpfulness, correctness,
coherence, conciseness). Preferences come from a prompted pairwise oracle (Qwen3-32B, one
attribute rubric per comparison, swap-averaged over presentation order); no reward model is fitted.
Base and reference are meta-llama/Llama-3.1-8B-Instruct.
All rules in this panel share prompts, samples and judged comparisons, and differ only in the per-objective weight applied inside the same update. Every rule was carried to the same four stages.
Win rate against the reference on 500 held-out prompts, averaged over two evaluation judges
(Phi-4 and Llama-3.3-70B), neither of which supplied the training preferences:
worst objective 0.369, average 0.410. Single seed.
- Downloads last month
- 12
Model tree for promotion/nbpo-helpsteer2-maxmin-stage4
Base model
meta-llama/Llama-3.1-8B