Llama-3.1-8B-PROSPER-baseline

PROSPER / MaxEntBW baseline, reproduced under the NBPO training setup. It selects each prompt's worst objective rather than bargaining across all four, and loses individual rationality on three of them.

Trained from meta-llama/Llama-3.1-8B-Instruct, which is also the reference policy (\mu) and the initialisation (\pi_0). Four objectives are scored on UltraFeedback prompts -- helpfulness, honesty, instruction following, truthfulness -- by a prompted Qwen3-32B preference oracle, each pair queried in both presentation orders and swap-averaged. Every arm in this release shares one pair set, one optimizer and a 300-step budget, and differs only in how the four objectives are aggregated, so a comparison between two of them isolates the aggregation rule.

Held-out surplus over the reference on the 657-prompt panel (population scale):

objective PROSPER
instruction following -0.0349
truthfulness -0.0508
honesty -0.0458
helpfulness +0.0676
minimum -0.0508

Decoding used for every reported benchmark: temperature 0.7, top-p 0.9, at most 1024 new tokens, seed 42. Generations for all arms are at promotion/nbpo-benchmark-generations so the comparison can be re-judged without re-decoding.

Built with Llama. Use is subject to the Llama 3.1 Community License.

Downloads last month
11
Safetensors
Model size
8B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for promotion/Llama-3.1-8B-PROSPER-baseline

Finetuned
(3042)
this model