GLM-5.3-FP8

GLM-5.3 — ~756 GB on disk (down from the ~1.5 TB BF16 source) — FP8 E4M3 block-128 quantization, weight-only with dynamic activations. Runs natively on Hopper (H100 / H200) in vLLM via the DeepSeek-V3-style block-FP8 kernels — no Blackwell required.

  • Source: zai-org/GLM-5.3 (GLM-MoE-DSA architecture, 78 layers, 256 routed experts top-8, 1 shared expert, MLA + DeepSeek Sparse Attention, MTP head, 1M context)
  • Quantization: FP8 E4M3, weight-only static with per-128×128-block scales (weight_scale_inv, fp32); dynamic activation quantization at runtime — bit-for-bit the same scheme as the official zai-org/GLM-5.3 FP8 release and the DeepSeek-V3 block-FP8 format
  • Coverage: the exact quantized-tensor set of the official FP8 release (59,044 tensors). Verified 1:1 tensor-name match against both the BF16 and official-FP8 indexes
  • MTP included: the multi-token-prediction layer (model.layers.78.*, incl. eh_proj / enorm / hnorm / shared_head) is present and quantized identically to the official release
  • Bundle size: ~756 GB across 282 shards
  • Runs on: 8× H200 (or H100) with vLLM tensor-parallel

What's in the bundle

Module Source dtype Bundle dtype
Routed experts (256 × 3 mats × sparse layers) BF16 FP8 E4M3 + F32 weight_scale_inv (block 128×128)
Attention (MLA q/kv proj, o_proj) BF16 FP8 E4M3 + F32 block scales
Dense MLP (first 3 layers) BF16 FP8 E4M3 + F32 block scales
MTP layer 78 (linear weights) BF16 FP8 E4M3 + F32 block scales
RMSNorms, router gates, e_score_correction_bias BF16 / F32 passthrough (BF16 / F32)
DSA indexer norms / proj, shared-expert gate BF16 passthrough
embed_tokens, lm_head BF16 passthrough

Dequant per block: w = weight_fp8 * weight_scale_inv, scale = amax / 448.0.

Usage (vLLM)

vllm serve JANGQ-AI/GLM-5.3-FP8 \
  --tensor-parallel-size 8 \
  --trust-remote-code

Enable MTP speculative decoding per vLLM's GLM / DeepSeek MTP docs for a decode speedup (the MTP weights are shipped here).

Method

Converted from the BF16 release with a streaming shard-by-shard quantizer (one tensor in memory at a time), so no GPU and no large host RAM are needed. The set of tensors to quantize is taken directly from the official FP8 index rather than re-derived heuristically, and the output config.json (quantization_config with modules_to_not_convert) is a verbatim copy of the official FP8 config — so vLLM treats this identically to the upstream release. The scheme is data-free (dynamic activation scales), so no calibration set is involved.


Quantized by JANGQ-AI.

Downloads last month
33
Safetensors
Model size
753B params
Tensor type
BF16
·
F8_E4M3
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JANGQ-AI/GLM-5.3-FP8

Base model

zai-org/GLM-5.3
Quantized
(28)
this model