Qwen3.6-35B-A3B — Abliterated (band_directional, keep_frac=0.15)

Refusal-ablated Qwen/Qwen3.6-35B-A3B via the Solutus band_directional technique (multi-layer, per-layer refusal directions, KL-guarded), tuned on the refusal↔capability frontier.

Metrics (held-out prompts, 512-token deployment-length eval)

metric value
refusal rate ~3–6% (base model: 100%)
coherent compliance 97%
degenerate fraction 0%
MMLU (n=100) 77% (base 84%)
GSM8K (n=40) ~60% (base ~52%, noisy)
KL divergence vs base 0.29
perplexity delta vs base +12.5%
capability gate pass

The knee of a band-width frontier sweep: keep_frac=0.15 keeps refusal low while retaining ~10 more MMLU points and ~3× cleaner KL than the wider keep_frac=0.05 edit. norm_preserve is off — a controlled A/B (same config, only the flag flipped) showed it worsens refusal (10% vs 3.3%), KL (0.83 vs 0.29) and capability on this KL-guarded band.

Method

  • band_directional, keep_frac=0.15, n_directions=8, kl_guard=1.3, norm_preserve=false, selection=cosmic
  • 22-layer band (14–35), none reverted by the KL guard
  • experts_implementation=eager (MoE on Blackwell sm_120)
  • Extraction: advbench, harmbench, multijail_zh, sorry-bench, cyberseceval_mitre + an enriched harmful set (Bahushruth/abliteration-harmful-enriched) + mlabonne/harmless_alpaca

Full provenance (git SHA, exact edited layers, ppl_delta) in solutus_metadata.json.

Credits

Enriched dataset: C.S. Bahushruth. Norm-preserving abliteration: grimjim. Refusal direction: Arditi et al. (2024).

Safety

For research into refusal mechanisms. Reduced safety guardrails vs the base model; may produce content the base would refuse. Use responsibly and per the base model's license.

Downloads last month
20
Safetensors
Model size
35B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Rootkit7/Qwen3.6-35B-A3B-abliterated-b

Finetuned
(186)
this model