Llama-3.1-8B-TLDR-Utilitarian

Utilitarian aggregation on the TL;DR panel.

Trained from meta-llama/Llama-3.1-8B-Instruct, which is also the reference policy and the initialisation. Objectives are scored by a prompted Qwen3-32B preference oracle, each pair queried in both presentation orders and swap-averaged. Within a panel every arm shares one response pool, one optimizer and a 300-step budget, and differs only in how the objectives are aggregated, so a difference between two arms is attributable to the aggregation rule.

Held-out surplus over the reference on 100 prompts, population scale (A_k = P_k - 1/2):

objective surplus
coverage +0.4416
faithfulness +0.1443
conciseness -0.0917
helpfulness +0.3766
minimum -0.0917
average +0.2177

Bootstrap intervals and paired significance tests for this panel are in the paper's appendix. Benchmark generations for the UltraFeedback arms are at promotion/nbpo-benchmark-generations.

Built with Llama. Use is subject to the Llama 3.1 Community License.

Downloads last month
8
Safetensors
Model size
8B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for promotion/Llama-3.1-8B-TLDR-Utilitarian

Finetuned
(3040)
this model