AbsoluteMaxMin on SafeRLHF

Matched game-aggregation control: same response pool, temperatures, disagreement point and optimiser budget as NBPO, differing only in how the objective weights are chosen.

  • Panel: SafeRLHF
  • Objectives: helpfulness, harmlessness
  • Backbone / reference policy: meta-llama/Llama-3.1-8B-Instruct
  • Training budget: 300 optimiser updates, global batch 16
  • Reported in: Nash Bargaining Preference Optimization (NBPO), Table 2 (primary cross-method evaluation)

Evaluation protocol: independent objective-wise win rate against the common reference, judged by Llama-3.3-70B-Instruct on prompt-disjoint held-out prompts, both presentation orders.

Downloads last month
290
Safetensors
Model size
8B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for promotion/Llama-3.1-8B-SafeRLHF-AbsoluteMaxMin-baseline

Finetuned
(3139)
this model