Llama-3.1-8B-PROSPER-baseline
PROSPER / MaxEntBW baseline, reproduced under the NBPO training setup. It selects each prompt's worst objective rather than bargaining across all four, and loses individual rationality on three of them.
Trained from meta-llama/Llama-3.1-8B-Instruct, which is also the reference policy (\mu) and the
initialisation (\pi_0). Four objectives are scored on UltraFeedback prompts -- helpfulness, honesty,
instruction following, truthfulness -- by a prompted Qwen3-32B preference oracle, each pair queried in
both presentation orders and swap-averaged. Every arm in this release shares one pair set, one optimizer
and a 300-step budget, and differs only in how the four objectives are aggregated, so a comparison
between two of them isolates the aggregation rule.
Held-out surplus over the reference on the 657-prompt panel (population scale):
| objective | PROSPER |
|---|---|
| instruction following | -0.0349 |
| truthfulness | -0.0508 |
| honesty | -0.0458 |
| helpfulness | +0.0676 |
| minimum | -0.0508 |
Decoding used for every reported benchmark: temperature 0.7, top-p 0.9, at most 1024 new tokens, seed 42.
Generations for all arms are at
promotion/nbpo-benchmark-generations
so the comparison can be re-judged without re-decoding.
Built with Llama. Use is subject to the Llama 3.1 Community License.
- Downloads last month
- 11
Model tree for promotion/Llama-3.1-8B-PROSPER-baseline
Base model
meta-llama/Llama-3.1-8B