YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

KAT‑Coder‑V2.5‑Dev‑MTP (int4 + MTP, vLLM‑ready)

Quantization MTP vLLM


This is a quantized, vLLM‑optimised version of KAT‑Coder‑V2.5‑Dev.
The model was post‑trained with int4 AutoRound (W4A16) and enhanced with a Multi‑Token Prediction (MTP) head from Qwen3.6‑35B‑A3B to enable speculative decoding in vLLM. It is designed for high‑throughput agentic coding workloads on hardware with 128 GiB unified memory (e.g. NVIDIA GB10).

All benchmark scores reported in the KAT‑Coder‑V2.5 technical paper refer to the full BF16 model.
This int4 quantisation has not been re‑evaluated on standard benchmarks. It is expected to closely match the original quality, but individual results may vary.


Model Overview

Property Value
Base model KAT‑Coder‑V2.5‑Dev (35 B total, 3 B active MoE)
Quantisation int4 AutoRound (W4A16), group size 128
MTP head Added from Qwen3.6‑35B‑A3B (Apache 2.0) for speculative decoding
Calibration data OpenCode Instruct (512 samples, sequence length 2048)
Quantisation tool Spark Auto Round – an optimised fork of Intel® AutoRound
Deployment target vLLM ≥ 0.19.0 with FlashInfer on GB10‑class hardware
Hugging Face Hub slopops/KAT-Coder-V2.5-Dev-MTP-int4-AutoRound-SAR

Key Features

  • Built on a highly efficient agentic coder – the original KAT‑Coder‑V2.5‑Dev was RL‑tuned to produce concise traces with fewer tool calls and reasoning tokens compared to its base model.
  • Speculative decoding with MTP – the added MTP head enables up to 3‑token lookahead. On GB10 hardware we observe throughput around 65–77 tokens/second under vLLM.
  • Reduced abnormal outputs – the original model reported tool‑label errors dropping from 9.34 % to 0.28 %, and single‑turn repetition nearly eliminated. These improvements are expected to carry over to the quantised version, though your mileage may vary.
  • Memory‑efficient int4 format – fits comfortably within 128 GiB unified memory, leaving headroom for large batch sizes and long contexts.

Deployment (vLLM)

This model is intended for use with vLLM only.
We recommend using a recent official image (≥ 0.19.0) with FlashInfer support.

docker run --rm --gpus all --net=host --ipc=host \
  vllm/vllm-openai:latest \
  vllm serve slopops/KAT-Coder-V2.5-Dev-MTP-int4-AutoRound-SAR \
    --port 8001 \
    --host 0.0.0.0 \
    --max-model-len 262144 \
    --gpu-memory-utilization 0.55 \
    --max-num-batched-tokens 16384 \
    --max-num-seqs 8 \
    --attention-backend flashinfer \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
    --tool-call-parser qwen3_coder \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --chat-template-kwargs '{"preserve_thinking":true}' \
    --generation-config '{"temperature":0.6,"top_p":0.95,"top_k":-1,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'

The server will be available at http://localhost:8001/v1.

Flag Reference

Flag Purpose
--attention-backend flashinfer Fast attention kernel, recommended for GB10
--speculative-config Activates MTP with 3 speculative tokens
--enable-auto-tool-choice + --tool-call-parser Required for agentic tool‑use
--chat-template-kwargs '{"preserve_thinking":true}' Keeps thinking traces from previous turns
--generation-config Safe defaults for deterministic outputs

Quantisation Details

  • Spark Auto Round – a GB10‑tuned fork of Intel® AutoRound that respects a 96 GiB memory budget during calibration.
  • Stateful resume – interrupted runs can be restarted from the last completed block.
  • ASAQ (Adaptive Sensitivity‑Aware Quantisation) – sensitive layers are identified via cosine similarity, PSNR, and router Jaccard similarity and kept in FP16 to preserve accuracy.
  • The MTP head was added after quantisation as a separate adapter; the core architecture remains the standard Qwen3.6‑35B‑A3B. vLLM handles speculative decoding via the --speculative-config flag.

Why This Model Exists

  • The original KAT‑Coder‑V2.5‑Dev was trained with hierarchical rewards that penalised excessive tool calls, empty blocks, and repetition – leading to more concise and reliable agentic traces.
  • The int4 quantisation makes it practical to run on hardware with 128 GiB unified memory without a separate GPU.
  • Adding the MTP head brings speculative decoding, which is especially beneficial for agentic loops where the model often emits multiple tokens in one step.

Performance claims (throughput, token savings, error reduction) refer to the original BF16 model as reported in the technical paper, or to our internal tests on a GB10 system. They are not guaranteed for every deployment. Always validate with your own workload.


Citation

If you use this quantised release, please cite the original KAT‑Coder work:

@misc{katcoder_v25_2026,
  title={{KAT-Coder-V2.5 Technical Report}},
  author={{KwaiKAT Team}},
  year={2026},
  month={July},
  eprint={2607.05471},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
  url={https://arxiv.org/pdf/2607.05471}
}

Disclaimer

  • Language‑model only – vision/multimodal components are not included.
  • All benchmark scores cited from the KAT‑Coder‑V2.5 paper are for the full BF16 model; this int4 version has not been re‑evaluated.
  • Throughput figures are observed on a specific hardware configuration (GB10 with 128 GiB unified memory) and may differ in other environments.
  • This model is provided “as‑is” for research and development purposes. Always test thoroughly before production use.
Downloads last month
-
Safetensors
Model size
1B params
Tensor type
I32
·
BF16
·
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for slopops/KAT-Coder-V2.5-Dev-MTP-int4-AutoRound-SAR