Llama-3.1-8B-HTMNPO-helpfulness
Single-objective corner of the weight simplex: all weight on helpfulness. Included to test what a degenerate vertex does to the objectives it ignores.
Trained from meta-llama/Llama-3.1-8B-Instruct, which is also the reference policy and the initialisation. Four objectives are scored on UltraFeedback prompts by a prompted Qwen3-32B preference oracle. Every arm in this release shares one pair set, one optimizer and the same budget, and differs only in how the objectives are aggregated.
Objective-wise held-out surplus over the reference (population scale, 100 prompts, prompted Qwen3-32B
oracle, swap-averaged over both presentation orders):
| objective | surplus |
|---|---|
| instruction following | -0.0571 |
| truthfulness | -0.0728 |
| honesty | -0.0668 |
| helpfulness | +0.0500 |
| minimum | -0.0728 |
Generations for every arm in the paper are at
promotion/nbpo-benchmark-generations.
Built with Llama. Use is subject to the Llama 3.1 Community License.
- Downloads last month
- -
Model tree for promotion/Llama-3.1-8B-HTMNPO-helpfulness
Base model
meta-llama/Llama-3.1-8B