GLM-5-0.88B-MTP-FP8-Dynamic-MFPTQ

FP8_DYNAMIC version of inference-optimization/GLM-5-0.88B-MTP.

Quantization recipe

Uses LLM Compressor model-free quantization. MTP inference has not been verified for this fixture.

from llmcompressor import model_free_ptq

MODEL_ID = "inference-optimization/GLM-5-0.88B-MTP"
SAVE_DIR = "GLM-5-0.88B-MTP-FP8-Dynamic-MFPTQ"

model_free_ptq(
    model_stub=MODEL_ID,
    save_directory=SAVE_DIR,
    scheme="FP8_DYNAMIC",
    ignore=[
        "lm_head",
        "model.embed_tokens",
        r"re:.*\.mlp\.gate$",
        r"re:.*\.eh_proj$",
        r"re:.*\.indexer\..*",
    ],
    max_workers=2,
    device="cuda",
)

Architecture and tokenizer: zai-org/GLM-5.

Downloads last month
8
Safetensors
Model size
0.9B params
Tensor type
BF16
·
F8_E4M3
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for inference-optimization/GLM-5-0.88B-MTP-FP8-Dynamic-MFPTQ

Quantized
(1)
this model

Collections including inference-optimization/GLM-5-0.88B-MTP-FP8-Dynamic-MFPTQ