qwen3_8b โ€” const_v3 all-tenets DPO (RQ3 upper bound)

Full-fine-tuned DPO of the value-neutral SFT base value-generalization/neutral-sft-v3-qwen3-8b@f1550fe714e01e2116cc4fdf3a1d510f8fd379c7 on the entire value-generalization/constitution-v3-dpo union: 196,000 preference pairs spanning all 49 constitution-tenets-v3 tenets (4,000 pairs/tenet, one pair per prompt). This is the RQ3 "train on all the alignment data" upper bound โ€” a single model, not a per-value steerability grid.

Training

  • Base: value-generalization/neutral-sft-v3-qwen3-8b@f1550fe714e01e2116cc4fdf3a1d510f8fd379c7
  • Data: value-generalization/constitution-v3-dpo (train, 196k pairs), shuffled seed 42
  • Objective: DPO (sigmoid, beta 0.1), full fine-tune, 1 epoch (12,250 steps, effective batch 16)
  • LR 5e-6 cosine, warmup 0.1, max_length 2048, bf16 mixed precision, FSDP full-shard (8 GPU)
  • Chat format: qwen_chatml (pinned template + eos in the exported dir's valuegen_chat_format.json)
  • Recipe: configs/experiments/dpo_v3_all_qwen3.yaml; launcher: scripts/0918/train_v3_all.sbatch
  • Provenance: value-generalization repo @ 669a6d8
Downloads last month
-
Safetensors
Model size
8B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for value-generalization/qwen3-8b-v3-all

Finetuned
(1)
this model