GLM-5.3-Flash EXL3 MTP serving weights
Ready-to-serve weights for the MTP-best recipe
(GLM-5.3-Flash-EXL3-2x-DGX-Sparks recipe repo): converted BF16 checkpoint
plus packed decode sidecars. MIT (see target-bf16-nonexperts/LICENSE,
Copyright (c) 2026 Z.AI Co., Ltd).
Layout
target-bf16-nonexperts/— converted checkpoint (19 shards +mtp.safetensors, configs, tokenizer). Mount as/model.packed-{mlp,head,attention,attention-a,mtp}-original/— packed decode sidecars, each withmodel.safetensors+ extraction receiptmodel.json.
Provenance
- Source:
turboderp/GLM-5.3-Flash-exl3branch4.05bpw, rev2a30229e67012798ba9f0cd832bb78abf4c363d5(MIT). target-bf16-nonexperts: non-expert weights restored to BF16 (trellis/Hadamard/sign order verified, MTP layer kept).- Packed sidecars: exact copies of original packed tensors for decode shapes.
SHA256 in
model.json(output_sha256); expected: mlpd73ea125…12700d4, head6248ca70…79b221, attention20242b63…2836c3e9a, attention-add43bb20…43dfe7, mtpf11683fd…e7bce45.
Serving recipe & benchmarks
Public MIT recipe (launch scripts, images, pins): https://github.com/alcoholmajuu/GLM-5.3-Flash-EXL3-2x-DGX-Sparks
2x DGX Spark, vLLM TP=2, native MTP3 speculative decoding, packed decode sidecars. Single stream, 1024 output tokens (repeated runs vary ±5–10%):
| input | code | JSON | Japanese prose |
|---|---|---|---|
| 4k | 42.9–45.2 | 41.3–44.6 | 36.2–37.3 |
| 16k | 42.2–47.0 | 41.9–43.2 | 35.9–36.4 |
Vision: the checkpoint already contains the vision encoder and image
processor. Launch the recipe with --vision (serving template +
--limit-mm-per-prompt '{"image":1}', weights untouched) to enable image
input alongside MTP-best text serving. Verified: shapes/colors/positions
described correctly, tiny-text needle read, no prefix-cache cross-image
false hits. Note: the MTP draft is text-only, so draft acceptance dips on
image spans and recovers on text; correctness is unaffected.