Kimi-K3 DFlash2 draft (lookahead-8, MLA + SWA)

Private draft-model artifact for speculative decoding on moonshotai/Kimi-K3. Serve with TokenSpeed (--speculative-algorithm DFLASH).

Provenance

field value
training run kimi-k3-dflash2-la8-full-v2-nofcnorm
checkpoint iter_0008001 (step 8000) — also best_checkpointed_iteration
TorchSpec source commit e6c34e321d3885c7abafecb026a0295e762d3e12
cluster NVIDIA TSD GB300 Slurm, 4 inference nodes (2x TP8) + 2 training nodes
fc_norm false (intended; the earlier e95a72e run had it unintentionally on)

Training-side eval at this checkpoint

metric value
eval/simulated_acc_len 1.7784578371886477
eval/avg_acc 0.4962460398674011
eval/avg_loss 2.306995391845703

Architecture

  • DFlash2DraftModel, 6 layers, hidden 7168, bf16, 103 tensors
  • block_size: 8 (lookahead-8), attention_mode: mla
  • target_layer_ids: [19, 37, 54, 66, 78, 90] (6 aux hidden states)
  • layer_types: 5x sliding_attention (window 4096) + 1x full_attention
  • yarn RoPE, rope_theta: 50000.0, factor: 32.0

Note: the validated serving path retains the full draft MLA KV cache in the unified full-attention storage group. Per-layer sliding-window metadata exists but the validated MLA backend does not consume it, so runs to date prove full-context MLA compute, not exact sliding_attention compute.

Export file hashes (sha256)

config.json        3d846e0b26714bf662b5d816cc27f52ec3d6b8e0d93f087f2dcc3296b42284cf
model.safetensors  fb54da2791a133473494a53ac8b5d93771a7a55de788ef924b3861179fc28a1f
Downloads last month
-
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support