Llama-3.1-8B-HTMNPO-helpfulness

Single-objective corner of the weight simplex: all weight on helpfulness. Included to test what a degenerate vertex does to the objectives it ignores.

Trained from meta-llama/Llama-3.1-8B-Instruct, which is also the reference policy and the initialisation. Four objectives are scored on UltraFeedback prompts by a prompted Qwen3-32B preference oracle. Every arm in this release shares one pair set, one optimizer and the same budget, and differs only in how the objectives are aggregated.

Objective-wise held-out surplus over the reference (population scale, 100 prompts, prompted Qwen3-32B oracle, swap-averaged over both presentation orders):

objective surplus
instruction following -0.0571
truthfulness -0.0728
honesty -0.0668
helpfulness +0.0500
minimum -0.0728

Generations for every arm in the paper are at promotion/nbpo-benchmark-generations.

Built with Llama. Use is subject to the Llama 3.1 Community License.

Downloads last month
-
Safetensors
Model size
8B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for promotion/Llama-3.1-8B-HTMNPO-helpfulness

Finetuned
(2977)
this model