Llama-3.1-8B-AbsoluteMaxmin-baseline
Absolute-maxmin baseline: all weight on the objective with the smallest raw game value rather than the smallest surplus over the reference.
Trained from meta-llama/Llama-3.1-8B-Instruct, which is also the reference policy and the
initialisation. Four objectives are scored on UltraFeedback prompts by a prompted Qwen3-32B preference
oracle, each pair queried in both presentation orders and swap-averaged. Every arm in this release shares
one pair set, one optimizer and a 300-step budget, and differs only in how the objectives are aggregated.
Held-out surplus over the reference on the 657-prompt panel (population scale):
| objective | surplus |
|---|---|
| instruction following | -0.0434 |
| truthfulness | -0.0624 |
| honesty | -0.0512 |
| helpfulness | +0.0594 |
| minimum | -0.0624 |
For comparison, the bargaining solution reaches a minimum of +0.0391 on the same panel:
promotion/Llama-3.1-8B-NBPO-600step.
Generations for every arm are at
promotion/nbpo-benchmark-generations.
Built with Llama. Use is subject to the Llama 3.1 Community License.
- Downloads last month
- -
Model tree for promotion/Llama-3.1-8B-AbsoluteMaxmin-baseline
Base model
meta-llama/Llama-3.1-8B