YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
KAT‑Coder‑V2.5‑Dev‑MTP (int4 + MTP, vLLM‑ready)
This is a quantized, vLLM‑optimised version of KAT‑Coder‑V2.5‑Dev.
The model was post‑trained with int4 AutoRound (W4A16) and enhanced with a Multi‑Token Prediction (MTP) head from Qwen3.6‑35B‑A3B to enable speculative decoding in vLLM. It is designed for high‑throughput agentic coding workloads on hardware with 128 GiB unified memory (e.g. NVIDIA GB10).
All benchmark scores reported in the KAT‑Coder‑V2.5 technical paper refer to the full BF16 model.
This int4 quantisation has not been re‑evaluated on standard benchmarks. It is expected to closely match the original quality, but individual results may vary.
Model Overview
| Property | Value |
|---|---|
| Base model | KAT‑Coder‑V2.5‑Dev (35 B total, 3 B active MoE) |
| Quantisation | int4 AutoRound (W4A16), group size 128 |
| MTP head | Added from Qwen3.6‑35B‑A3B (Apache 2.0) for speculative decoding |
| Calibration data | OpenCode Instruct (512 samples, sequence length 2048) |
| Quantisation tool | Spark Auto Round – an optimised fork of Intel® AutoRound |
| Deployment target | vLLM ≥ 0.19.0 with FlashInfer on GB10‑class hardware |
| Hugging Face Hub | slopops/KAT-Coder-V2.5-Dev-MTP-int4-AutoRound-SAR |
Key Features
- Built on a highly efficient agentic coder – the original KAT‑Coder‑V2.5‑Dev was RL‑tuned to produce concise traces with fewer tool calls and reasoning tokens compared to its base model.
- Speculative decoding with MTP – the added MTP head enables up to 3‑token lookahead. On GB10 hardware we observe throughput around 65–77 tokens/second under vLLM.
- Reduced abnormal outputs – the original model reported tool‑label errors dropping from 9.34 % to 0.28 %, and single‑turn repetition nearly eliminated. These improvements are expected to carry over to the quantised version, though your mileage may vary.
- Memory‑efficient int4 format – fits comfortably within 128 GiB unified memory, leaving headroom for large batch sizes and long contexts.
Deployment (vLLM)
This model is intended for use with vLLM only.
We recommend using a recent official image (≥ 0.19.0) with FlashInfer support.
docker run --rm --gpus all --net=host --ipc=host \
vllm/vllm-openai:latest \
vllm serve slopops/KAT-Coder-V2.5-Dev-MTP-int4-AutoRound-SAR \
--port 8001 \
--host 0.0.0.0 \
--max-model-len 262144 \
--gpu-memory-utilization 0.55 \
--max-num-batched-tokens 16384 \
--max-num-seqs 8 \
--attention-backend flashinfer \
--enable-prefix-caching \
--enable-chunked-prefill \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--chat-template-kwargs '{"preserve_thinking":true}' \
--generation-config '{"temperature":0.6,"top_p":0.95,"top_k":-1,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
The server will be available at
http://localhost:8001/v1.
Flag Reference
| Flag | Purpose |
|---|---|
--attention-backend flashinfer |
Fast attention kernel, recommended for GB10 |
--speculative-config |
Activates MTP with 3 speculative tokens |
--enable-auto-tool-choice + --tool-call-parser |
Required for agentic tool‑use |
--chat-template-kwargs '{"preserve_thinking":true}' |
Keeps thinking traces from previous turns |
--generation-config |
Safe defaults for deterministic outputs |
Quantisation Details
- Spark Auto Round – a GB10‑tuned fork of Intel® AutoRound that respects a 96 GiB memory budget during calibration.
- Stateful resume – interrupted runs can be restarted from the last completed block.
- ASAQ (Adaptive Sensitivity‑Aware Quantisation) – sensitive layers are identified via cosine similarity, PSNR, and router Jaccard similarity and kept in FP16 to preserve accuracy.
- The MTP head was added after quantisation as a separate adapter; the core architecture remains the standard Qwen3.6‑35B‑A3B. vLLM handles speculative decoding via the
--speculative-configflag.
Why This Model Exists
- The original KAT‑Coder‑V2.5‑Dev was trained with hierarchical rewards that penalised excessive tool calls, empty blocks, and repetition – leading to more concise and reliable agentic traces.
- The int4 quantisation makes it practical to run on hardware with 128 GiB unified memory without a separate GPU.
- Adding the MTP head brings speculative decoding, which is especially beneficial for agentic loops where the model often emits multiple tokens in one step.
Performance claims (throughput, token savings, error reduction) refer to the original BF16 model as reported in the technical paper, or to our internal tests on a GB10 system. They are not guaranteed for every deployment. Always validate with your own workload.
Citation
If you use this quantised release, please cite the original KAT‑Coder work:
@misc{katcoder_v25_2026,
title={{KAT-Coder-V2.5 Technical Report}},
author={{KwaiKAT Team}},
year={2026},
month={July},
eprint={2607.05471},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/pdf/2607.05471}
}
Disclaimer
- Language‑model only – vision/multimodal components are not included.
- All benchmark scores cited from the KAT‑Coder‑V2.5 paper are for the full BF16 model; this int4 version has not been re‑evaluated.
- Throughput figures are observed on a specific hardware configuration (GB10 with 128 GiB unified memory) and may differ in other environments.
- This model is provided “as‑is” for research and development purposes. Always test thoroughly before production use.
- Downloads last month
- -