GLM-5.3-Flash EXL3 MTP serving weights

Ready-to-serve weights for the MTP-best recipe (GLM-5.3-Flash-EXL3-2x-DGX-Sparks recipe repo): converted BF16 checkpoint plus packed decode sidecars. MIT (see target-bf16-nonexperts/LICENSE, Copyright (c) 2026 Z.AI Co., Ltd).

Layout

  • target-bf16-nonexperts/ — converted checkpoint (19 shards + mtp.safetensors, configs, tokenizer). Mount as /model.
  • packed-{mlp,head,attention,attention-a,mtp}-original/ — packed decode sidecars, each with model.safetensors + extraction receipt model.json.

Provenance

  • Source: turboderp/GLM-5.3-Flash-exl3 branch 4.05bpw, rev 2a30229e67012798ba9f0cd832bb78abf4c363d5 (MIT).
  • target-bf16-nonexperts: non-expert weights restored to BF16 (trellis/Hadamard/sign order verified, MTP layer kept).
  • Packed sidecars: exact copies of original packed tensors for decode shapes. SHA256 in model.json (output_sha256); expected: mlp d73ea125…12700d4, head 6248ca70…79b221, attention 20242b63…2836c3e9a, attention-a dd43bb20…43dfe7, mtp f11683fd…e7bce45.

Serving recipe & benchmarks

Public MIT recipe (launch scripts, images, pins): https://github.com/alcoholmajuu/GLM-5.3-Flash-EXL3-2x-DGX-Sparks

2x DGX Spark, vLLM TP=2, native MTP3 speculative decoding, packed decode sidecars. Single stream, 1024 output tokens (repeated runs vary ±5–10%):

input code JSON Japanese prose
4k 42.9–45.2 41.3–44.6 36.2–37.3
16k 42.2–47.0 41.9–43.2 35.9–36.4

Vision: the checkpoint already contains the vision encoder and image processor. Launch the recipe with --vision (serving template + --limit-mm-per-prompt '{"image":1}', weights untouched) to enable image input alongside MTP-best text serving. Verified: shapes/colors/positions described correctly, tiny-text needle read, no prefix-cache cross-image false hits. Note: the MTP draft is text-only, so draft acceptance dips on image spans and recovers on text; correctness is unaffected.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support