GLM-5.3 DFlash2 drafters (GGUF)

GGUF conversions of incoai/GLM-5.3-DFlash2 — the DFlash2 block-diffusion drafter for GLM-5.3 753B (zai-org/GLM-5.3).

File Size
GLM-5.3-DFlash2-BF16.gguf 4.59 GiB — lossless from source
GLM-5.3-DFlash2-Q8_0.gguf 2.44 GiB
-md GLM-5.3-DFlash2-Q8_0.gguf --spec-type draft-dflash --spec-draft-n-max 7

block_size is 8, so 7 is the ceiling — llama.cpp clamps anything higher. Lower values can win when the target's experts are partly CPU-resident, since each verification step is then more expensive; sweep it for your setup.

Works with any GLM-5.3 GGUF — DFlash consumes the target's hidden states (hidden_size 6144, layers [6,20,34,48,62,76]), not its quantisation. Not compatible with GLM-5.3-Flash (hidden 4096).

Downloads last month
211
GGUF
Model size
2B params
Architecture
dflash
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for emwesoft/GLM-5.3-DFlash2-GGUF

Base model

zai-org/GLM-5.3
Quantized
(2)
this model