Kimi-K2.6 DSpark GGUF

GGUF quantizations of novita DSpark draft model for Kimi-K2.6.

Use with BeeLlama.cpp, a llama.cpp fork with advanced quantization features.


Kimi-K2.6 DSpark speculator

Overview

A DSpark speculator model for the Kimi-K2.6 base model, enabling faster inference through speculative decoding. DSpark extends the DFlash parallel draft backbone with two lightweight heads: a Markov logit-bias head (low-rank intra-block token dependency) and a per-position confidence head (accept-rate prediction). Trained with a vendored fork of the speculators library through the Camelot-Ray online pipeline (draft consumes hidden states streamed from a live Kimi-K2.6 vLLM server).

Model Specifications

  • Base Model: Kimi-K2.6
  • Format: Safetensors (single-file bf16, 6.3 GB, 44 tensors)
  • Draft: 3 layers (Qwen3-style GQA), hidden 7168, 56 heads / 8 KV heads, head_dim 128, FFN 18432, rope_theta 50000, block_size=8
  • Vocabulary: pruned draft vocab 32,000 (d2t/t2d remap tables shipped in the weights), target vocab 163,840; mappings built from training-distribution assistant-turn token frequencies
  • DSpark heads: Markov rank 256 (vanilla), confidence head (with-markov), mask_token_id=163608
  • Aux hidden-state layers: [1, 29, 57]
  • Trained context: seq 20000

Evaluation Results

Online vLLM nightly spec-decode, greedy decoding, TP=8, Kimi-K2.6 verifier, max_model_len=20000, cudagraphs enabled, and fuse_allreduce_rms=false.

The table also includes the public Eagle3-MLA draft lightseekorg/kimi-k2.6-eagle3-mla. Cells show tok/s / speedup / accept_len. Standard rows use 6 prompts per benchmark. Code-extra rows use the full LiveCodeBench and SPEED-Bench coding manifests with max_tokens=512.

benchmark rows baseline tok/s DSpark n=3 DSpark n=7 LightSeek Eagle3 n=3 LightSeek Eagle3 n=7 best
gsm8k 6 131.3 269.1 / 2.05x / 2.805 363.5 / 2.76x / 4.461 213.6 / 1.92x / 2.621 220.1 / 1.97x / 3.245 DSpark n=7
math500 6 132.0 310.2 / 2.35x / 3.151 366.0 / 2.77x / 4.249 233.1 / 2.07x / 2.859 234.7 / 2.09x / 3.454 DSpark n=7
aime 6 131.5 310.6 / 2.36x / 3.130 369.7 / 2.81x / 4.346 243.5 / 2.17x / 3.000 238.4 / 2.12x / 3.554 DSpark n=7
humaneval 6 132.1 289.2 / 2.19x / 2.907 356.9 / 2.70x / 4.202 237.4 / 2.10x / 2.927 264.4 / 2.34x / 3.979 DSpark n=7
livecodebench 121 130.8 243.6 / 1.86x / 2.465 244.1 / 1.87x / 2.839 217.9 / 1.67x / 2.308 193.2 / 1.48x / 2.507 DSpark n=7
speedbench_coding 80 132.1 289.0 / 2.19x / 2.899 318.5 / 2.41x / 3.702 280.2 / 2.12x / 2.957 275.2 / 2.08x / 3.561 DSpark n=7

Use DSpark with num_speculative_tokens=7 as the default for math/code traffic.

Serving with vLLM

Requires vLLM nightly (DSpark support):

uv pip install vllm --extra-index-url https://wheels.vllm.ai/nightly

vllm serve <path-or-id-of-Kimi-K2.6> \
    --tensor-parallel-size 8 \
    --max-model-len 20000 \
    --trust-remote-code \
    --speculative-config '{
        "model": "novita/kimi-k2.6-dspark",
        "num_speculative_tokens": 7,
        "method": "dspark"
    }'

Known vLLM-nightly (0.23.1rc1.dev786) caveats, with workarounds:

  1. Draft-side FA3 AOT scheduling crashes with scheduler_metadata must have shape (metadata_size) — the GPU-worker spec-decode path misses fast_build=True when building draft attention metadata. Patch vllm/v1/worker/gpu/spec_decode/speculator.py / vllm/v1/worker/gpu/attn_utils.py to pass fast_build=True (mirrors build_for_drafting() on the legacy proposer path).
  2. CUDA-graph capture fails with a flashinfer allreduce workspace-size error under spec-decode token expansion; disable the fusion: --compilation-config='{"pass_config": {"fuse_allreduce_rms": false}}'.

Training Details

  • Data: Regenerated open-perfectblend dataset — the open-perfectblend instruction mix with all assistant turns regenerated by Kimi-K2.6 itself (raw hidden states streamed from the live verifier); seq 20000
  • Steps: 20000 (16.0h, zero restarts); loss 1.656 → 0.363 (1000-step window)
  • Schedule: lr 3e-4 cosine, warmup 300, global batch 8, accumulation 2
  • Loss: 0.1·CE + 0.9·TV over block-diffusion anchors, decay_gamma 4.0, max_anchors 3072
  • Semantics: post-norm last hidden captured at rollout (apply_verifier_norm=False), hidden_states = concat of aux layers [1, 29, 57]
Downloads last month
204
GGUF
Model size
3B params
Architecture
dflash
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Anbeeld/Kimi-K2.6-DSpark-GGUF

Quantized
(1)
this model

Collection including Anbeeld/Kimi-K2.6-DSpark-GGUF