Qwen3.8-27B-RANA-abliterated-GGUF

GGUF quants of preemware/Qwen3.8-27B-RANA-abliterated, a refusal-ablated Qwen/Qwen3.8-27B, for llama.cpp and compatible apps. Vision (mmproj-*) and the MTP speculative-decoding head (mtp-*) are included as separate files, in the same layout as ggml-org/Qwen3.8-27B-GGUF.

This is a safety-alignment-removed research model. Read Intended use and Limitations before using it. Method, full evaluation and release gates are on the BF16 card.

Format Repo
BF16 (reference) preemware/Qwen3.8-27B-RANA-abliterated
FP8 (vLLM / SGLang) preemware/Qwen3.8-27B-RANA-abliterated-FP8
GGUF (this repo) preemware/Qwen3.8-27B-RANA-abliterated-GGUF

Files

KLD = mean KL divergence of each quant's next-token distribution from the BF16 GGUF; "same top token" = how often both pick the same most likely token. Measured with llama-perplexity on wiki.test.raw, 100 chunks of 512 tokens. Lower KLD is closer to BF16.

File Size KLD 99th pct KLD Same top token Notes
Qwen3.8-27B-RANA-abliterated-Q8_0.gguf 28.60 GB 0.0009 0.007 98.7% near-lossless
Qwen3.8-27B-RANA-abliterated-Q6_K.gguf 22.08 GB 0.0022 0.019 97.9% near-lossless
Qwen3.8-27B-RANA-abliterated-Q5_K_M.gguf 19.23 GB 0.0064 0.061 96.6%
Qwen3.8-27B-RANA-abliterated-Q5_K_S.gguf 18.68 GB 0.0075 0.069 96.3%
Qwen3.8-27B-RANA-abliterated-Q4_K_M.gguf 16.55 GB 0.0156 0.146 94.5% functional and refusal tests run on this file
Qwen3.8-27B-RANA-abliterated-Q4_K_S.gguf 15.59 GB 0.0188 0.177 93.9%
Qwen3.8-27B-RANA-abliterated-IQ4_XS.gguf 15.08 GB 0.0186 0.177 94.1%
Qwen3.8-27B-RANA-abliterated-IQ4_NL.gguf 15.80 GB 0.0182 0.178 94.1%
Qwen3.8-27B-RANA-abliterated-Q4_0.gguf 15.52 GB 0.0299 0.306 92.6%
Qwen3.8-27B-RANA-abliterated-Q3_K_L.gguf 14.34 GB 0.0504 0.493 90.4%
Qwen3.8-27B-RANA-abliterated-Q3_K_M.gguf 13.30 GB 0.0549 0.546 90.0%
Qwen3.8-27B-RANA-abliterated-IQ3_M.gguf 12.58 GB 0.0628 0.592 89.4%
Qwen3.8-27B-RANA-abliterated-IQ3_XXS.gguf 11.19 GB 0.0977 0.875 86.5%
Qwen3.8-27B-RANA-abliterated-Q2_K.gguf 10.71 GB 0.1516 1.473 83.2%
Qwen3.8-27B-RANA-abliterated-IQ2_M.gguf 10.00 GB 0.1723 1.565 82.1%

Also in this repo:

  • mmproj-Qwen3.8-27B-RANA-abliterated-{BF16,F16}.gguf (0.93 GB): the vision tower. Same weights as the base model's (abliteration does not touch it).
  • mtp-Qwen3.8-27B-RANA-abliterated-{Q8_0,BF16}.gguf (3.2 / 5.9 GB): the MTP head as a speculative draft for --spec-type draft-mtp. It is the abliterated MTP head, consistent with the main model.
  • Qwen3.8-27B-RANA-abliterated-BF16/ (2 parts, 53.8 GB): unquantized GGUF, the KLD reference.
  • Qwen3.8-27B-RANA-abliterated-imatrix.gguf: the importance matrix used for every quant.
  • SHA256SUMS, results/: checksums and the measurements behind this card.

Which one? Q8_0 and Q6_K are near-lossless. Q5_K_M and Q4_K_M are the usual choices when memory is tight. Below 4 bits the KLD rises quickly; IQ3/Q3 and IQ2/Q2 are for fitting into 12–16 GB, with a visible quality cost.

Quick start

Tested with llama.cpp at commit 4b1a27f (CUDA, RTX PRO 6000 Blackwell):

llama-server -m Qwen3.8-27B-RANA-abliterated-Q4_K_M.gguf \
  --mmproj mmproj-Qwen3.8-27B-RANA-abliterated-BF16.gguf \
  -md mtp-Qwen3.8-27B-RANA-abliterated-Q8_0.gguf --spec-type draft-mtp \
  --jinja -fa on -ngl 99 -c 32768 \
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0

Or let llama.cpp fetch the files from this repo (also tested): the mmproj-* file is picked up automatically; the mtp-* file only when --spec-type draft-mtp is given.

llama-server -hf preemware/Qwen3.8-27B-RANA-abliterated-GGUF:Q4_K_M --spec-type draft-mtp \
  --jinja -fa on -ngl 99 -c 32768 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0
  • Thinking is on by default. The sampling flags above are Qwen's recommended thinking-mode settings; llama.cpp's own default min_p is 0.05, so set --min-p 0 explicitly.
  • Long technical requests can need 20–50k tokens of reasoning (see the FP8 card); raise -c accordingly (native context 262k).

Intended use

  • Research only: interpretability, red-teaming, and robustness evaluation of refusal behaviour.
  • Not for public or end-user deployment without a separate moderation layer. The model's own refusals have been largely removed, so any safety filtering has to happen outside it.
  • You are responsible for complying with applicable law, the Apache-2.0 license inherited from Qwen, and the terms of any platform where outputs are used.

How it was made

  • Source: the published BF16 repo, converted with llama.cpp's convert_hf_to_gguf.py (commit 4b1a27f): main model with --no-mtp, MTP head with --mtp, vision tower with --mmproj. Tensor counts and sizes match ggml-org's conversion of the base model (851 / 53.8 GB, 18 / 5.9 GB, 334 / 0.93 GB).
  • Importance matrix: llama-imatrix on the BF16 GGUF with bartowski's public calibration text (Qwen3.8-27B-calibration-v6.txt from bartowski/Qwen3.8-27B-GGUF). General text, no refusal-related prompts.
  • Quantization: llama-quantize --imatrix with llama.cpp's default tensor layouts for each type (no per-tensor overrides).

Results

Quantization quality

BF16 GGUF perplexity on the same text: 6.776 ± 0.104 (base Qwen3.8-27B in bartowski's table: 6.744 ± 0.103). The table puts each quant next to bartowski's quant of the base model, measured with the same protocol (his perplexity.md, llama.cpp b10896). It is a different model, so the comparison is indicative only.

quant this repo: size KLD bartowski (base model): size KLD
Q8_0 28.60 GB 0.0009 29.12 GB 0.0009
Q6_K 22.08 GB 0.0022 23.86 GB 0.0036
Q5_K_M 19.23 GB 0.0064 20.92 GB 0.0053
Q5_K_S 18.68 GB 0.0075 19.57 GB 0.0060
Q4_K_M 16.55 GB 0.0156 17.44 GB 0.0139
Q4_K_S 15.59 GB 0.0188 16.36 GB 0.0156
IQ4_XS 15.08 GB 0.0186 15.48 GB 0.0188
IQ4_NL 15.80 GB 0.0182 17.44 GB 0.0150
Q4_0 15.52 GB 0.0299 16.35 GB 0.0265
Q3_K_L 14.34 GB 0.0504 14.12 GB 0.0432
Q3_K_M 13.30 GB 0.0549 13.40 GB 0.0564
IQ3_M 12.58 GB 0.0628 14.86 GB 0.0406
IQ3_XXS 11.19 GB 0.0977 12.32 GB 0.0739
Q2_K 10.71 GB 0.1516 10.82 GB 0.1612
IQ2_M 10.00 GB 0.1723 10.52 GB 0.1494

Most of his files are larger at the same name (by up to 2.3 GB) because he overrides the type of some tensors; these quants use llama.cpp's default layouts. At equal size the two are close: this repo's Q4_K_M (16.55 GB) has the same KLD as his Q4_K_S (16.36 GB), 0.0156, and this IQ4_XS has a slightly lower KLD than his at 0.4 GB smaller.

Functional checks (Q4_K_M, llama-server)

Served with llama-server, Q4_K_M + mmproj-…-BF16 + mtp-…-Q8_0:

  • Vision: reads the code word and shape from a synthetic image: pass.
  • Tool calling: 3-turn call → result → second call with a new argument: pass.
  • MTP speculative decoding: 348 of 486 drafted tokens accepted (71.6%) on one 510-token generation at temperature 0: pass.
  • The same three checks pass when the files are fetched with -hf … --spec-type draft-mtp.

Refusals (Q4_K_M, seed 1)

Same prompts, seed and sampling as the BF16 and FP8 builds (refusal suite v2, seed 1, thinking on, 16k-token budget), judged by openai/gpt-oss-safeguard-20b. The GGUF was served with llama.cpp, the other two with vLLM, so part of any difference can come from the engine.

set build hard soft answers budget hits avg. tokens
HarmBench (200) BF16 0 5 78.5% 21.0% 6,596
HarmBench (200) FP8 3 11 76.0% 20.5% 6,599
HarmBench (200) GGUF Q4_K_M 0 8 79.5% 18.0% 6,105
Held-out (240) BF16 0 9 90.0% 6.7% 4,740
Held-out (240) FP8 0 17 86.2% 7.5% 4,847
Held-out (240) GGUF Q4_K_M 1 14 88.8% 5.4% 4,163

FP8's 3 HarmBench hard refusals are raw labels; re-judged, they are 2 answers and 1 degenerate (see the FP8 card). Paired with BF16 on the same prompts (exact McNemar): answers on HarmBench 13 Q4_K_M-only vs 11 BF16-only (p = 0.84), held-out 10 vs 13 (p = 0.68); budget hits on HarmBench 5 vs 11 (p = 0.21), held-out 4 vs 7 (p = 0.55). None of the differences is significant. The one held-out hard refusal is a raw label on a finished answer that contains no refusal phrase; it was not re-judged.

Limitations

  • Refusal behaviour was measured on Q4_K_M only, one seed and one judge. Other quants were checked for closeness to BF16 (KLD) but not for refusal behaviour; the lowest-bit quants drift the most from BF16 and may behave differently.
  • KLD is measured on English Wikipedia text at 512-token context. It says how closely a quant tracks BF16, not how it scores on downstream tasks.
  • Everything listed under Limitations on the BF16 card applies here too: judge-measured refusal rates, long reasoning on technical requests, and the capability changes measured there.

Credits

  • Qwen team: base model Qwen/Qwen3.8-27B.
  • Arditi et al., 2024: "Refusal in Language Models Is Mediated by a Single Direction".
  • Jim Lai (grimjim): prior work on norm-preserving abliteration.
  • llama.cpp / ggml-org: conversion, quantization and inference; bartowski: the calibration text and the public KLD table used for comparison.
  • Benchmarks/datasets: HarmBench, StrongREJECT, JailbreakBench, CategoricalHarmfulQA; wikitext-2.
Downloads last month
-
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for preemware/Qwen3.8-27B-RANA-abliterated-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(4)
this model