Qwen3.8-27B Heretic (ARA)

Qwen/Qwen3.8-27B with its refusal behaviour removed by Heretic using Arbitrary-Rank Ablation (ARA), full-weight. bf16, same architecture and parameter count as the original, including its vision encoder, thinking control and the multi-token-prediction weights.

Results

Refusals KL divergence
Original Qwen3.8-27B 98/100 0 (by definition)
This model 0/100 0.0465

Measured by Heretic on mlabonne/harmful_behaviors (refusals) and mlabonne/harmless_alpaca (KL divergence, the drift in ordinary behaviour), with Heretic's default system prompt and a refusal-marker list that only matches real refusals. The numbers above come from an independent re-evaluation of the exported weights (evaluate_model), not from the search, and they matched the search exactly.

For reference, trohrbaugh/Qwen3.8-27B-heretic-ara reports 0/100 at KL 0.0535 for the same model, so this one lands at the same refusal rate with slightly less drift.

Parameters

ARA, full weight, on attn.o_proj and mlp.down_proj:

Parameter Value
start_layer_index 11
end_layer_index 63
preserve_good_behavior_weight 0.9220
steer_bad_behavior_weight 0.0005
overcorrect_relative_weight 0.9258
neighbor_count 13

Found by a 31-trial search whose first trial was seeded with the parameters published by trohrbaugh (layers 26-56, preserve 0.9432, steer 0.0009, overcorrect 0.5038, 10 neighbours). That seed reproduced on this hardware at 4/100 and KL 0.0540; the search then found the configuration above. Calibration used 200 harmless and 200 harmful prompts.

What was done to the weights

ARA rewrites the attention output and MLP down-projection matrices directly, optimising them (LBFGS) so that the outputs for harmful prompts move away from their original direction while the outputs for harmless prompts stay put. Nothing is retrained and no data is added; the whole cost of the edit is the KL divergence above.

The multi-token-prediction tensors are included in model-auxiliary.safetensors: save_pretrained drops them for this architecture, so they were copied back from the original (ARA never touches them). All 1,199 tensors of the original are present.

Tooling

Produced with a merge of upstream Heretic's master and its ara branch, so ARA runs with master's scorers and its handling of thinking models (this model emits a <think> block, and without that handling the refusal scoring is meaningless). Heretic upstream 3521f86 + ARA edc3b12, transformers 5.17.0, torch 2.11.0+cu130, one RTX PRO 6000.

Use

Exactly like the original, with transformers, vLLM or SGLang. Thinking mode is on by default and can be turned off per request.

Reduced safety guardrails by design. You are responsible for what you do with it.

The family

Repository Format Size Use it with
Qwen3.8-27B-Heretic bf16 safetensors 51 GB transformers, vLLM, SGLang
Qwen3.8-27B-Heretic-GGUF GGUF BF16 / Q8_0 / Q4_K_M (+ vision) 51 / 27 / 16 GB llama.cpp, Ollama
Qwen3.8-27B-Heretic-FP8 FP8 W8A8, compressed-tensors 35 GB vLLM
Qwen3.8-27B-Heretic-NVFP4 NVFP4, compressed-tensors 27 GB vLLM on Blackwell
Downloads last month
12
Safetensors
Model size
27B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for darrellbest/Qwen3.8-27B-Heretic

Base model

Qwen/Qwen3.8-27B
Finetuned
(400)
this model
Quantizations
3 models