Llama-3.1-8B-TLDR-HTMNPO-coverage
Single-objective corner on the TL;DR panel: all weight on coverage.
Trained from meta-llama/Llama-3.1-8B-Instruct, which is also the reference policy and the
initialisation. Objectives are scored by a prompted Qwen3-32B preference oracle, each pair queried in
both presentation orders and swap-averaged. Within a panel every arm shares one response pool, one
optimizer and a 300-step budget, and differs only in how the objectives are aggregated, so a difference
between two arms is attributable to the aggregation rule.
Held-out surplus over the reference on 100 prompts, population scale (A_k = P_k - 1/2):
| objective | surplus |
|---|---|
| coverage | +0.1011 |
| faithfulness | -0.0196 |
| conciseness | -0.0145 |
| helpfulness | +0.0905 |
| minimum | -0.0196 |
| average | +0.0394 |
Bootstrap intervals and paired significance tests for this panel are in the paper's appendix. Benchmark
generations for the UltraFeedback arms are at
promotion/nbpo-benchmark-generations.
Built with Llama. Use is subject to the Llama 3.1 Community License.
- Downloads last month
- 1
Model tree for promotion/Llama-3.1-8B-TLDR-HTMNPO-coverage
Base model
meta-llama/Llama-3.1-8B