qwen3_8b โ const_v3 all-tenets DPO (RQ3 upper bound)
Full-fine-tuned DPO of the value-neutral SFT base value-generalization/neutral-sft-v3-qwen3-8b@f1550fe714e01e2116cc4fdf3a1d510f8fd379c7 on the entire
value-generalization/constitution-v3-dpo union: 196,000 preference pairs
spanning all 49 constitution-tenets-v3 tenets (4,000 pairs/tenet, one pair per
prompt). This is the RQ3 "train on all the alignment data" upper bound โ a
single model, not a per-value steerability grid.
Training
- Base:
value-generalization/neutral-sft-v3-qwen3-8b@f1550fe714e01e2116cc4fdf3a1d510f8fd379c7 - Data:
value-generalization/constitution-v3-dpo(train, 196k pairs), shuffled seed 42 - Objective: DPO (sigmoid, beta 0.1), full fine-tune, 1 epoch (12,250 steps, effective batch 16)
- LR 5e-6 cosine, warmup 0.1, max_length 2048, bf16 mixed precision, FSDP full-shard (8 GPU)
- Chat format:
qwen_chatml(pinned template + eos in the exported dir'svaluegen_chat_format.json) - Recipe:
configs/experiments/dpo_v3_all_qwen3.yaml; launcher:scripts/0918/train_v3_all.sbatch - Provenance: value-generalization repo @
669a6d8
- Downloads last month
- -
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support
Model tree for value-generalization/qwen3-8b-v3-all
Base model
value-generalization/neutral-sft-v3-qwen3-8b