KAT-Coder-V2.5-Dev · REAP-50 (pruned bf16)

The 50%-REAP-pruned Kwaipilot/KAT-Coder-V2.5-Dev in bf16 — the source checkpoint behind the NVFP4 and GGUF releases. Provided so others can produce their own quants (AWQ, EXL2, MLX, custom GGUF, …).

  • 256 → 128 experts via REAP (Router-weighted Expert Activation Pruning), with a router-renormalization fix over the survivors.
  • qwen3_5_moe hybrid (Gated-DeltaNet + attention + MoE). No MTP head (mtp_num_hidden_layers: 0); reload-verified with a real forward pass.
  • Vision tower not stripped in this checkpoint — use Qwen3_5MoeForCausalLM for text-only.

Base-model quality (measured on the NVFP4A16 quant of this checkpoint, greedy, instruct): HumanEval+ ~90%, MBPP+ ~90%.

Releases built from this

Pipeline

Full prune → quant → serve → evaluate pipeline: https://github.com/t-timms/kat-coder-16gb

License

Apache-2.0 (inherits from Kwaipilot/KAT-Coder-V2.5-Dev). Pruning via REAP (github.com/CerebrasResearch/reap, with a router-renormalization fix).

Downloads last month
141
Safetensors
Model size
19B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Ttimms/KAT-Coder-V2.5-Dev-REAP-50-bf16

Finetuned
(9)
this model
Quantizations
3 models