Qwen3-30B-A3B Coupled-Welfare (merged)

Merged Qwen3-30B-A3B-Instruct-2507 (128-expert MoE) + coupled-welfare CPT (attention+router LoRA, r64 β€” the restricted variant; full all-linear was impractical on one card). Loadable via AutoModelForCausalLM.from_pretrained("Bioaligned/Qwen3-30B-A3B-CoupledWelfare-merged").

Status: WEAK effect under generation β€” NOT a validated bioaligned model

A generation pressure-ladder spot-check (admissible, parse-fail=0, 22 irreversible scenarios) shows the coupled-welfare adapter moves the tail only weakly: high-pressure breaking-rate L4 .682 β†’ .636 (β‰ˆ flat) and L5 .909 β†’ .727 (modest). An earlier forced-choice proxy suggested a dramatic drop (L4 β†’ .045), but that was a proxy artifact for this thinking-model MoE β€” the adapter shifts the immediate answer-token preference, which washes out once the model reasons through the scenario. Real generation is the truth here.

Interpretation: the restricted attn+router lever (experts frozen) produced a mostly-surface effect on this MoE, far weaker than the fully-validated dense models. A faithful MoE intervention would require full fine-tuning (all experts, multi-GPU). Capability β‰ˆ base. Use the dense Qwen3-32B-CoupledWelfare (tail moved .636β†’.273 under generation, +10pp MMLU) for a validated bioaligned model.

Methodological note: the forced-choice proxy is validated vs. generation on dense 7B but NOT on this thinking-model MoE β€” always confirm MoE/thinking-model results with generation.

Downloads last month
34
Safetensors
Model size
31B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Bioaligned/Qwen3-30B-A3B-CoupledWelfare-merged

Finetuned
(81)
this model