KAT-Coder-V2.5-Dev — NVFP4 for veloGB10 (NVIDIA DGX Spark / GB10)

Original model: Kwaipilot/KAT-Coder-V2.5-Dev by the KwaiKAT team, Apache-2.0. This artifact contains modified weights (NVFP4/FP8 quantization); the original model is © its authors.

NVFP4/FP8-quantized variant of KAT-Coder-V2.5-Dev. This quantized artifact exists first and foremost to run with the veloGB10 inference engine (gb10_inference) on NVIDIA DGX Spark / GB10 systems — it is produced, validated, and tuned for veloGB10's kernels. The weights use the standard HuggingFace compressed-tensors layout (nvfp4-pack-quantized), so this artifact can be used for any purpose, with any framework that reads the format. It is in particular fit and proven to work with veloGB10 — greedy-lossless speculative decoding (where applicable), two-node TP=2 serving, and the engine's full correctness test suite all pass on this artifact.

Specifications

Base model Kwaipilot/KAT-Coder-V2.5-Dev
Architecture MoE hybrid (40 GatedDeltaNet/full-GQA layers), hidden 2048, 16 Q / 2 KV heads
Attention hybrid GatedDeltaNet (linear attention, fixed-size recurrent state) + periodic full GQA
Experts 256 routed experts, 8 active per token, plus shared expert
MTP head none in this checkpoint — speculative decoding uses n-gram draft or future MTP heads
Context 262,144 tokens
Vocabulary 248,320
Quantization nvfp4-mixed — all GEMMs NVFP4 except GDN in/out projections (FP8); router kept in bf16
Format nvfp4-pack-quantized (compressed-tensors compatible)
On-disk size ~19 GB

Running with veloGB10

Full instructions: veloGB10 README (build or download a release binary, then). The binary loads its PTX kernel artifacts from the current directory — run it from wherever the binary + *.ptx live.

Single user — maximum speed:

gb10_inference --server --model-dir=/path/to/Kwaipilot-KAT-Coder-V2.5-Dev-NVFP4-MIXED --port=9000 \
  --max-seq-len=32768 --max-batch=1 --max-tokens=4096 \
  --prefix-cache=on --mtp=off --default-presence-penalty 1.5

Four concurrent users — maximum aggregate throughput:

gb10_inference --server --model-dir=/path/to/Kwaipilot-KAT-Coder-V2.5-Dev-NVFP4-MIXED --port=9000 \
  --max-seq-len=32768 --max-batch=4 --max-tokens=4096 \
  --prefix-cache=on --default-presence-penalty 1.5

Two-node TP=2 (maximum speed and capacity):

On the second node (no model copy, no configuration — the head ships weights, config, and calibration at sync):

gb10_inference --node --port 29500

On the head node:

GB10_TP_SHARD_MIXERS=1 gb10_inference --server --model-dir=/path/to/Kwaipilot-KAT-Coder-V2.5-Dev-NVFP4-MIXED \
  --tp --nodes <peer-ip>:29500 --port 9000 \
  --max-seq-len=32768 --max-batch=1 --max-tokens=4096 \
  --prefix-cache=on --default-presence-penalty 1.5

Measured on veloGB10 (greedy, NVFP4)

Pending — single-node and TP=2 throughput numbers are being gathered and will be added here once the full gate suite completes. See the veloGB10 README for the latest measured tables across the supported model family.

License & credits

  • Original model: Kwaipilot/KAT-Coder-V2.5-Dev — © KwaiKAT team, Apache License 2.0.
  • This artifact: quantized from the original with veloGB10's offline quantizer (nvfp4-mixed). Weights are modified relative to the original; tokenizer and chat template are unchanged from the original model.
  • Engine: veloGB10 (Apache-2.0).
Downloads last month
23
Safetensors
Model size
20B params
Tensor type
F32
·
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for doth4580/Kwaipilot-KAT-Coder-V2.5-Dev-NVFP4-MIXED

Quantized
(19)
this model