GLM-5.3 753B — NVFP4 GGUF (no MTP)

Native GGUF of incoai/GLM-5.3-NVFP4, the vendor's ModelOpt NVFP4 repack of zai-org/GLM-5.3 (753B, glm-dsa, 78 blocks, hidden 6144, 256 experts). The NVFP4 weights are kept as NVFP4 — this is not a requantisation.

445 GB, 10 shards. block_count 78, 1947 tensors, 225 NVFP4 tensors.

Related repos

With MTP head emwesoft/GLM-5.3-NVFP4-MTP-GGUF — adds block 78, enables --spec-type draft-mtp
DFlash2 drafters emwesoft/GLM-5.3-DFlash2-GGUF — speculative decoding for this model

This variant has no MTP head, so --spec-type draft-mtp is unavailable. Use the DFlash2 drafters for speculative decoding, or the MTP repo above.

Engine requirements

llama.cpp with glm-dsa + GGML_TYPE_NVFP4. Three fixes are not yet upstream:

  1. jinja numeric attribute access (obj.0) — GLM-5.3's chat template uses m.content.0.output. Without it the template throws, caps_get() swallows it, supports_tool_calls reports false, and every tool call comes back as plain text.
  2. glm-dsa layer-input exposure — DFlash needs res->t_layer_inp[il]; without it attaching a drafter aborts on the first decode with GGML_ASSERT(t_layer_inp[il] != nullptr).
  3. Whitespace tolerance before </tool_call> — a stray newline makes the streaming parser recognise a tool call then lose it, aborting from compute_diffs.

Measured throughput

2x RTX PRO 6000 Blackwell + 4x RTX 3090 + 251 GB RAM, 400K context, -t 36 -tb 40, experts partly CPU-resident (the weights do not fit in 288 GB of VRAM):

config acceptance decode
MTP head, n-max 3 74.9% (mean len 3.24) 9.3-12.2 tok/s
DFlash2 Q8_0, n-max 4 67.7% (mean len 3.69) 7.6-12.8 tok/s
no speculation - ~10 tok/s

Throughput is prompt-dependent because acceptance is. Threads matter: on a 24-core/48-thread CPU, -t 48 collapsed decode to 0.5 tok/s — the ggml threadpool busy-spins and starves the CUDA submission thread. Leave headroom.

Sampling

From generation_config.json: temperature 1.0, top_p 0.95. The template exposes low/high/max reasoning effort only; anything else becomes max, and thinking cannot be disabled.

Downloads last month
336
GGUF
Model size
743B params
Architecture
glm-dsa
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for emwesoft/GLM-5.3-NVFP4-GGUF

Base model

zai-org/GLM-5.3
Quantized
(2)
this model