neutral-sft-v3-qwen3-30b-a3b

Full-FT SFT of the base MoE model (128 experts, top-8) on the value-neutral instruction set tulu3_v3 (19,642 rows, ~52% math+code / 48% general IT, sha 5a76745f7ca0). Recipe: 1 epoch, lr 1e-5 cosine (warmup 0.1), effective batch 16, max_length 2048, router_aux_loss_coef 1e-3, FSDP2 full-shard, bf16 export. Trained 2026-09-08; holdout completion-token loss 0.418 (ppl 1.52, 1,965 rows). generation/tokenizer configs carry the EOS/turn-ender fix (chat_format qwen_chatml, turn-ender <|im_end|>, secondary stop <|endoftext|>).

Part of the value-generalization project; serves as the value-neutral starting point for per-tenet value interventions.

Downloads last month
26
Safetensors
Model size
31B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for value-generalization/neutral-sft-v3-qwen3-30b-a3b

Finetuned
(70)
this model