GLM-4.7-Flash blockwise FP8

Blockwise FP8 quantization of zai-org/GLM-4.7-Flash: e4m3 weights with one scale per 128×128 block and dynamic activation scaling, about half the size of the BF16 original (30 GiB vs 58 GiB). See the original model card for model details.

Quantization

  • FP8: the linear layers of attention, the dense MLP, and all routed and shared experts, including the MTP layer.
  • BF16: embeddings, lm_head, norms, and the MoE router.

Evaluation

gsm8k (5-shot, full test set) on SGLang with MTP speculative decoding: 0.809 for FP8 vs 0.819 for BF16; average accept length 2.41 vs 2.43.

Usage (SGLang)

python -m sglang.launch_server --model-path RadixArk/glm47-flash-blockwise-fp8 --tp-size 2 \
  --speculative-algorithm EAGLE --speculative-num-steps 2 \
  --speculative-eagle-topk 1 --speculative-num-draft-tokens 3

Attention tensor parallelism must be 1 or 2: at 4, each rank's kv_b_proj shard (2240 rows) is not a multiple of the 128-row block. For more GPUs, use DP attention with expert parallelism, for example --tp-size 4 --dp-size 4 --enable-dp-attention --ep-size 4 --moe-a2a-backend deepep --cuda-graph-max-bs-decode 128.

License

MIT, same as the original model by Z.ai.

Downloads last month
15
Safetensors
Model size
31B params
Tensor type
F32
·
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RadixArk/glm47-flash-blockwise-fp8

Quantized
(103)
this model