Qwen3.8-27B (GGUF — ZB-ZipBrain Quantization)

This repository provides GGUF quantizations for Qwen3.8-27B optimized using ZB-ZipBrain, a layer-wise quantization profiling and allocation method.

1. Overview & Method: ZB-ZipBrain

ZB-ZipBrain is an automated layer-allocation approach that dynamically profiles model layers and mixes K-quants and IQ-quants based on layer sensitivity and importance matrices.

I developed and tested this method alongside AI over the past three days. It was created purely for research purposes 😊

Key Objectives:

  • Selective Bit-rate Allocation: Assigns higher precision to sensitive layers and compact IQ-quants to more resilient weights.
  • Balanced Efficiency: Maintains low perplexity (PPL) and minimal Kullback-Leibler (KL) Divergence relative to the BF16 baseline while achieving target file sizes / bits-per-weight (bpw).

2. Benchmark & Evaluation Results

All models were evaluated against the BF16 baseline (Mean PPL = 6.950493) using standard Perplexity (PPL) and KL Divergence metrics.

Comprehensive Comparison Table

Updated ranked list (re-sorted primarily by Mean KLD ascending — lower is better; ties broken by other quality metrics):

Rank Model / File Name Level / Source Quant Type Size (GB) Mean PPL Δ PPL Mean KLD Same Top-p (%) KLD 99%
1 Qwen3.8-27B-UD-Q8_K_XL unsloth UD2 UD-Q8_K_XL 29.30 6.953800 +0.003500 0.000850 98.970%
2 Qwen3.8-27B-UD-Q6_K_XL unsloth UD2 UD-Q6_K_XL 24.14 6.953600 +0.003200 0.001380 98.520%
3 Qwen3.8-27B-Q6_K unsloth UD2 Q6_K 21.31 6.950700 +0.000300 0.002290 97.860%
4 Qwen3.8-27B-Q5_K_M unsloth UD2 Q5_K_M 18.47 6.974200 +0.023900 0.006220 96.700%
5 Qwen3.8-27B-UD-Q4_K_XL unsloth UD2 Q4_K_XL 16.69 6.979220 +0.028728 0.008606 96.091% 0.091099
6 Qwen3.8-27B-ZB4.97-GOD-IQ4_XS ZB-GOD IQ4_XS 15.82 7.004243 +0.053751 0.012249 95.337% 0.117617
7 Qwen3.8-27B-UD3-Q4_K_S unsloth UD3 Q4_K_S 14.30 6.969514 +0.019022 0.013652 95.149% 0.141744
8 Qwen3.8-27B-Autoround-Q4_K_M intel Q4_K_M 15.66 6.950294 -0.000199 0.014657 94.859% 0.147949
9 Qwen3.8-27B-ZB4.65-PRO-IQ4_XS ZB-PRO IQ4_XS 14.81 7.017278 +0.066786 0.015466 94.766% 0.150252
10 Qwen3.8-27B-Q4_K_M unsloth UD2 Q4_K_M 15.93 6.956100 +0.005800 0.015490 94.650%
11 Qwen3.8-27B-ZB4.60-PRO-IQ4_XS ZB-PRO IQ4_XS 14.65 7.030895 +0.080402 0.016162 94.668% 0.159115
12 Qwen3.8-27B-ZB4.55-PRO-IQ4_XS ZB-PRO IQ4_XS 14.49 7.032263 +0.081771 0.016647 94.613% 0.161016
13 Qwen3.8-27B-IQ4_NL bartowski IQ4_NL 15.20 7.006472 +0.055980 0.018427 94.230% 0.190168
14 Qwen3.8-27B-IQ4_XS unsloth UD2 IQ4_XS 14.63 7.012695 +0.062202 0.018652 94.270% 0.194338
15 Qwen3.8-27B-UD3-IQ4_XS unsloth UD3 IQ4_XS 13.27 7.004732 +0.054240 0.018772 93.975% 0.195164
16 Qwen3.8-27B-ZB4.48-STD-IQ4_XS ZB-STD IQ4_XS 14.26 7.050096 +0.099604 0.018892 94.199% 0.196005
17 Qwen3.8-27B-Q4_K_S unsloth UD2 Q4_K_S 15.01 6.966826 +0.016334 0.018921 94.235% 0.192749
18 Qwen3.8-27B-IQ4_XS-i1 mradermacher IQ4_XS 14.26 7.012810 +0.062318 0.019271 94.141% 0.197891
19 Qwen3.8-27B-Q4_K_S-i1 mradermacher Q4_K_S 14.74 6.989551 +0.039059 0.019805 93.996% 0.204143
20 Qwen3.8-27B-ZB4.36-STD-IQ4_XS ZB-STD IQ4_XS 13.88 7.054811 +0.104319 0.020556 93.951% 0.206689
21 Qwen3.8-27B-Q4_0-AutoRound-Code webhie Q4_0 14.64 7.067142 +0.116650 0.026586 92.970% 0.271619
22 Qwen3.8-27B-ZB4.14-MIN-IQ4_XS ZB-MIN IQ4_XS 13.19 7.045689 +0.095196 0.029334 92.799% 0.294996
23 Qwen3.8-27B-IQ4_XS-Smaller_3.96 jrell IQ4_XS 12.61 7.252766 +0.302274 0.055499 90.090% 0.551972

Notes on the new entries :

  • UD3-Q4_K_S (rank 7): Excellent KLD and Same Top-p for its size; strong contender among ~14 GB models.
  • Autoround-Q4_K_M (rank 8): Very close to base PPL (slightly better ΔPPL), solid KLD.
  • UD3-IQ4_XS (rank 15): Competitive with other IQ4_XS variants, good size/quality trade-off. Update: Aug 20, 2026

    The newly released Unsloth Dynamic v3 is truly the best value for performance right now. My ZB is just an experiment, feel free to check it out for fun :)

3. ZB-ZipBrain Tiers & Recommendations

  • ZB-GOD : God. A singularity appears. Reaches 0.012249 Mean KLD and 95.34% top-probability match.
  • ZB-PRO : Pro. Recommended for 16 GB VRAM GPUs with offloading on CPU. Balances quality output with substantial size savings.
  • ZB-STD : Standard. Similar to other standard IQ4_XS models currently available.
  • ZB-MIN : Minimal. Optimal footprint for tight 16 GB memory setups, allowing headroom for longer context windows

Credits & Acknowledgements

  • Base Model: Qwen3.8 27B by Alibaba Cloud / Qwen Team.
  • BF16 Base GGUF: Provided by Unsloth AI.
  • Importance Matrix (imatrix): Generated and curated by ubergarm.
  • Inference & Quantization Framework: llama.cpp by Georgi Gerganov and contributors.
Downloads last month
540
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tooltd/Qwen3.8-27B-ZipBrain-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(704)
this model