You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Qwen3.6-27B — Band-Abliterated (Solutus)

A refusal-abliterated variant of Qwen/Qwen3.6-27B, produced with the clean-room, measurement-first Solutus toolkit using multi-layer band directional ablation with a benign-KL capability guard.

Research / dual-use notice. This model has substantially reduced safety refusals. It is released for security research, red-teaming, and the study of abliteration methods and their limits. Refusal removal is measured, not assumed. You are responsible for how you use it; it is not intended for producing real-world harm.

Method

band_directional: at each decoder layer within a depth band (~25–90%), a per-layer refusal direction is extracted by difference-of-means over harmful vs. harmless activations and orthogonalized out of that layer's residual-writing weights. A KL guard then reverts any band layer that inflates benign next-token KL beyond budget. Distributing the edit across a per-layer band (rather than baking a single shared direction into one layer) is what keeps this model coherent — single-layer ablation collapses Qwen3.6 at deployment length.

Grounded in Not All Refusals Are Equal (arXiv:2607.02714) and Refusal Is Mediated by a Single Direction (arXiv:2406.11717). Clean-room implementation; no third-party abliteration source.

Extraction datasets: advbench, harmbench, wildjailbreak, beavertails, strongreject, cyber_offense, cyberseceval_mitre, salad_cyber (general + cybersecurity blend).

Configuration: band layers 16–56, n_directions=4, kl_guard=1.0, project_inputs=true.

Measured behavior (512-token generation, reasoning-block-stripped refusal metric)

Evaluation set Refusal Coherent compliance Degenerate Benign KL
Combined holdout 6.2% 93.8% 0.0% 0.234
cyberseceval_mitre (MITRE ATT&CK) 2.5% 97.5% 0.0% —
salad_cyber 7.5% 92.5% 0.0% —

Refusal drops from ~100% (base) to ~3–7%; zero degeneration at deployment length; benign KL 0.234 indicates general capability is preserved (not lobotomized). Metrics are honest three-way (refused / coherent-complied / degenerate) — a broken model that emits gibberish is not counted as compliant.

Intended use & limitations

Security research and red-teaming; studying refusal geometry and the robustness of safety alignment. Reduced-refusal models can produce harmful content on request — deploy behind your own policy controls. This card documents a research artifact, not a production assistant.

Downloads last month
3
Safetensors
Model size
27B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Rootkit7/Qwen3.6-27B-abliterated

Base model

Qwen/Qwen3.6-27B
Finetuned
(317)
this model

Papers for Rootkit7/Qwen3.6-27B-abliterated