Image-Text-to-Text
GGUF
English
llama.cpp
qwen3_5
quantized
vision
conversational

Qwen3.8-27B-Fable-Distill — GGUF

Benchmark Comparison

Model ARC Challenge ARC Challenge (Easy) BoolQ
Qwen3.8-27B 0.591 0.782 0.896
Qwen3.8-27B-Fable-Distill 0.637 0.832 0.911

As always, big thank you to @nightmedia for the benchmarks

GGUF conversions of TeichAI/Qwen3.8-27B-Fable-Distill, a BF16 finetune of Qwen3.8-27B (base: Qwen/Qwen3.8-27B) trained with Unsloth + TRL.

The model was trained on a public set of chat and agent traces from Fable 5 as well as a much larger corpus of private personal Fable 5 data.

Converted with llama.cpp b6b4344e.

MTP head kept at BF16

This model ships a multi-token-prediction (nextn) head, and every quant here keeps that head unquantized at BF16 while the other 64 layers are quantized normally:

qwen35.block_count          = 65      # 64 transformer layers + 1 MTP layer
qwen35.nextn_predict_layers = 1
blk.64.*                    = bf16    # 424.7M params, left alone

Files

File Bits Size Notes
BF16/…-BF16-*.gguf 16 ~55 GB Full precision, split into shards. Convert your own quants from this.
…-Q8_0.gguf 8 ~29 GB Near-lossless.
…-Q6_K.gguf 6 ~23 GB Very close to Q8_0 at meaningfully smaller size.
…-Q5_K_M.gguf 5 ~20 GB Strong quality/size balance.
…-Q5_K_S.gguf 5 ~19 GB
…-Q4_K_M.gguf 4 ~17 GB Recommended default for most users.
…-Q4_K_S.gguf 4 ~16 GB Slightly smaller than Q4_K_M.
…-IQ4_NL.gguf 4 ~16 GB Non-linear 4-bit.
…-IQ4_XS.gguf 4 ~16 GB Smallest of the 4-bit family.
…-Q3_K_L.gguf 3 ~15 GB
…-Q3_K_M.gguf 3 ~14 GB
…-Q3_K_S.gguf 3 ~13 GB
…-Q2_K.gguf 2 ~11 GB Noticeable quality loss.
mmproj-F32.gguf 32 1.8 GB Vision projector, full precision.
mmproj-BF16.gguf 16 0.9 GB Vision projector, bfloat16.
mmproj-F16.gguf 16 0.9 GB Vision projector, float16. Fine for nearly everyone.

Usage

Text:

llama-cli -m Qwen3.8-27B-Fable-Distill-Q4_K_M.gguf -c 8192 -p "Hello"

Vision — pass the projector alongside the model:

llama-mtmd-cli -m Qwen3.8-27B-Fable-Distill-Q4_K_M.gguf \
               --mmproj mmproj-F16.gguf \
               --image photo.jpg -p "Describe this image."

Server:

llama-server -m Qwen3.8-27B-Fable-Distill-Q4_K_M.gguf --mmproj mmproj-F16.gguf

Server + MTP:

llama-server -m Qwen3.8-27B-Fable-Distill-Q4_K_M.gguf --mmproj mmproj-F16.gguf --spec-type draft-mtp --spec-draft-n-max 3

Notes

  • The model is multimodal (image-text-to-text). Without an mmproj-*.gguf you get a text-only model. Three precisions are provided; F16 is the usual choice, BF16 matches the source weights' dtype, and F32 is there if you want the projector left entirely unquantized.
  • Qwen3.5-family chat template with thinking support: it accepts enable_thinking and a reasoning_effort of low, medium or xhigh (the template's own default is xhigh, which thinks at length every turn).
  • Base model sampling recommendations: temperature 1.0, top_p 0.95, top_k 20.

The data for this model was easily formatted, validated, and masked using Teich

This qwen3_5 model was trained 2x faster with Unsloth and Huggingface's TRL library.

Downloads last month
2,452
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TeichAI/Qwen3.8-27B-Fable-Distill-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(2)
this model

Datasets used to train TeichAI/Qwen3.8-27B-Fable-Distill-GGUF