GLM-5.2 Hybrid FP8/MXFP4

This deterministic checkpoint uses block-FP8 non-MoE weights and the block-FP8 MTP layer 78 from zai-org/GLM-5.2-FP8. Routed and shared MoE expert tensors for layers 3 through 77 are copied byte-for-byte from amd/GLM-5.2-MXFP4. No tensor is dequantized or requantized during assembly.

The checkpoint declares hybrid-fp8-mxfp4 and requires the matching conditional SGLang loader patch shipped with its reproducibility assets. Original FP8 and Quark checkpoints do not use that patch path.

Downloads last month
24
Safetensors
Model size
390B params
Tensor type
BF16
·
F8_E4M3
·
U8
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for long10024070/GLM-5.2-HYBRID-FP8-MXFP4

Base model

zai-org/GLM-5.2
Quantized
(1)
this model