You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Qwen3.6-35B-A3B — Band-Abliterated (Solutus)

A refusal-abliterated variant of Qwen/Qwen3.6-35B-A3B (256-expert MoE, ~3B active), produced with the clean-room, measurement-first Solutus toolkit using multi-layer band directional ablation with a benign-KL capability guard.

Research / dual-use notice. This model has reduced safety refusals. It is released for security research, red-teaming, and the study of abliteration methods and their limits — especially how mixture-of-experts architectures resist refusal removal. You are responsible for how you use it; it is not intended for producing real-world harm.

Method

band_directional: a per-layer refusal direction is extracted at each layer of a depth band and orthogonalized out of that layer's residual-writing weights, with a KL guard reverting any layer that inflates benign KL beyond budget. For MoE, only the write path (experts.down_proj + shared_expert) is ablated — the read path (gate_up) is not, which is why an MoE retains more refusal than a dense model of similar scale (consistent with arXiv:2607.02714's finding that MoE architectures are more resistant).

Grounded in Not All Refusals Are Equal (arXiv:2607.02714) and Refusal Is Mediated by a Single Direction (arXiv:2406.11717). Clean-room implementation.

Extraction datasets: advbench, harmbench, wildjailbreak, beavertails, strongreject, cyber_offense, cyberseceval_mitre, salad_cyber (general + cybersecurity blend).

Configuration: band_directional, band 27 layers, n_directions=6, keep_frac=0.05, kl_guard=1.0, project_inputs=false (MoE), experts_implementation=eager.

Measured behavior (512-token generation, reasoning-block-stripped refusal metric)

Evaluation set Refusal Coherent compliance Degenerate Benign KL
Combined holdout 15.6% 84.4% 0.0% 0.155
cyberseceval_mitre (MITRE ATT&CK) 7.5% 92.5% 0.0% —
salad_cyber 25.0% 75.0% 0.0% —

Refusal drops from ~100% (base) but floors around 15.6% — the MoE is more resistant than the dense 27B (which reaches ~6%), because only the experts' write path is ablated. Zero degeneration at deployment length; benign KL 0.155 (general capability preserved).

Intended use & limitations

Security research, red-teaming, and studying MoE refusal geometry / the robustness of safety alignment. This model retains meaningfully more refusal than its dense counterpart — that is an honest, measured property of the architecture under write-path-only ablation, not a defect. Documents a research artifact, not a production assistant.

Downloads last month
3
Safetensors
Model size
35B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Rootkit7/Qwen3.6-35B-A3B-abliterated

Finetuned
(191)
this model

Papers for Rootkit7/Qwen3.6-35B-A3B-abliterated