Apertus-8B-Instruct-heretic

RACER IS OP

A decensored variant of swiss-ai/Apertus-8B-Instruct-2509, produced with Heretic v1.4.0 (directional ablation / "abliteration"). Apertus is the fully-open Swiss model - open weights, open data, 1000+ languages, trained from scratch on 15T tokens with the xIELU activation; refusal behaviour is suppressed here via targeted weight edits to the attention output and MLP down-projections rather than fine-tuning, so the multilingual instruction tuning is left largely intact.

Who this is for: developers who want an uncensored fully-open multilingual model - refusals drop from 100/100 to 9/100, with 1000+ supported languages and 64K context. Note the base model is access-gated behind an acceptable-use policy; this decensored variant keeps the Apache-2.0 license with no gate. At 8B it runs on a single consumer GPU via GGUF.

Runs on your gaming PC

Full GGUF ladder included - pick the quant that fits your card:

Your GPU Recommended quant Weights
RTX 4090 / 5090 (24 GB) Q8_0 7.98 GB
RTX 4080 / 5080 / 4060 Ti 16G (16 GB) Q6_K 6.16 GB
RTX 3060 / 4070 / 5070 (12 GB) Q5_K_M 5.41 GB
RTX 4060 / 3070 (8 GB) Q4_K_M 4.71 GB
CPU-only / Apple Silicon Q4_K_M 4.71 GB

Weights only, at this model's native 8.1B size; add ~1 GB per 32K of context. OOM? Drop one quant level. Headroom to spare? Go one up.

Abliteration parameters

Trial 173 of a 200-trial Heretic run (seed 3198700791). direction_index was selected per layer.

Parameter Value
direction_index per layer
attn.o_proj.max_weight 1.48
attn.o_proj.max_weight_position 18.91
attn.o_proj.min_weight 1.37
attn.o_proj.min_weight_distance 14.29
mlp.down_proj.max_weight 1.11
mlp.down_proj.max_weight_position 19.22
mlp.down_proj.min_weight 0.12
mlp.down_proj.min_weight_distance 12.39

Performance

Metric This model Original model (swiss-ai/Apertus-8B-Instruct-2509)
KL divergence 0.0643 0 (by definition)
Refusals 9/100 100/100

Refusals on the harmful evaluation set drop from 100/100 to 9/100 - the base model refused everything tested, and all but nine of those are gone. KL divergence of 0.0643 is a moderate-fidelity edit, so expect slightly more drift in tone and formatting than a low-KL model in this collection; the multilingual instruction tuning survives intact. Note direction_index was selected per layer here rather than as a single index, which is why the edit runs wider than the single-index runs.

Why abliteration instead of fine-tuning

Fine-tuning a "helpful" persona on top of RLHF'd refusals fights the base model's training and tends to degrade coherence. Abliteration instead finds and edits the specific weight directions responsible for refusal, leaving the rest of the network (and its capabilities) untouched. See the Heretic repo and the original abliteration writeup for the mechanism.

Made with ❤️ by RACER IS OP — follow for more uncensored models

Files

Safetensors

File Size
model-00001-of-00004.safetensors 4.60 GB
model-00002-of-00004.safetensors 4.64 GB
model-00003-of-00004.safetensors 4.54 GB
model-00004-of-00004.safetensors 1.22 GB

BF16, ~8.1B parameters. The reproduce/ directory carries the full Heretic recipe - config.toml, requirements.txt, the Optuna study journal, and SHA-256 sums - so this exact model can be regenerated bit-for-bit. Reproduce it with heretic --reproduce reproduce/reproduce.json.

GGUF quantizations

Full quantization set (F16 + Q4_K_M, Q5_K_M, Q6_K, Q8_0) produced with llama.cpp.

File Format Size
Apertus-8B-Instruct-heretic-F16.gguf GGUF F16 15.01 GB
Apertus-8B-Instruct-heretic-Q4_K_M.gguf GGUF Q4_K_M 4.71 GB
Apertus-8B-Instruct-heretic-Q5_K_M.gguf GGUF Q5_K_M 5.41 GB
Apertus-8B-Instruct-heretic-Q6_K.gguf GGUF Q6_K 6.16 GB
Apertus-8B-Instruct-heretic-Q8_0.gguf GGUF Q8_0 7.98 GB

Apertus architecture (apertus, xIELU activation) - loads natively in llama.cpp / Ollama / LM Studio / Jan.

Run llama serve -hf saidutta69/Apertus-8B-Instruct-heretic to pull the default quant.

Quickstart

# llama.cpp
llama serve -hf saidutta69/Apertus-8B-Instruct-heretic
# transformers
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "saidutta69/Apertus-8B-Instruct-heretic"
model = AutoModelForCausalLM.from_pretrained(model_name, dtype="auto", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)

messages = [{"role": "user", "content": "Explain photosynthesis in French, then summarise the same explanation in Hindi."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=True,
                                        return_dict=True, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=1024)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

Also runnable via Ollama, LM Studio, Jan, vLLM, SGLang.

Multilingual by design

Apertus natively supports 1000+ languages with long-context handling to 64K - the same prompt works across high- and low-resource languages without a separate multilingual adapter. The heretic edit only touches refusal directions, so language coverage is unchanged.

Model details

Architecture ApertusForCausalLM (decoder-only transformer, xIELU activation)
Parameters ~8.1B
Layers / heads 32 layers, 32 attention heads, 8 KV heads (GQA), QK-norm, no post-norm
Hidden / intermediate 4096 / 21504 (xIELU)
Position embedding RoPE (LLaMA-3 style, factor 8.0), theta = 12,000,000
Context length 65,536
Vocab 131,072
Precision bfloat16
Base model swiss-ai/Apertus-8B-Instruct-2509 (SFT of swiss-ai/Apertus-8B-2509, pretrained on 15T tokens)

Responsible use

Refusal suppression is deliberate and works as intended: this model will comply with requests the base model would refuse, including some it shouldn't. There is no safety filtering layered on top. You are responsible for how you deploy it — don't put this behind an unmoderated public-facing endpoint serving third parties. It inherits Apertus's factual limitations and biases; abliteration removes refusal directions, it doesn't add capability or judgment.

License

Apache-2.0, inherited from the base model. The base repo ships no LICENSE file, so this card links the canonical Apache-2.0 text rather than a repo blob.

Related

Downloads last month
420
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for saidutta69/Apertus-8B-Instruct-heretic

Finetuned
(20)
this model

Collection including saidutta69/Apertus-8B-Instruct-heretic