GLM-5.3-Flash EXL3 suite

Suite status: pending

Public selective-EXL3 variants of zai-org/GLM-5.3-Flash-BF16, pinned at a6c167b62691b2bac901344b65cb651a70f53e43.

Target Repository Payload Structural Runtime Quality
3.0 bpw 3.0 pending pending pending pending
2.5 bpw 2.5 pending pending pending pending
2.0 bpw 2.0 pending pending pending pending

No variant in this table is released yet. Each repository is card-only until its complete weight snapshot, manifest, checksums, public state, runtime evidence, and measured quality status are independently audited. The Collection will be created and linked only after all required variants pass publication verification.

Only routed-expert gate/up/down projections are quantized. The accuracy-sensitive backbone and shared path remain BF16. The artifacts use a custom TP4 selective-EXL3 layout; stock Transformers compatibility is not claimed.

Credits: Z.AI for the MIT-licensed base model, ExLlamaV3/TurboDerp for EXL3, and Dione for the selective conversion workflow.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 0xSero/GLM-5.3-Flash-EXL3

Quantized
(11)
this model