Kimi-K3 DFlash2 draft (lookahead-8, MLA + SWA)
Private draft-model artifact for speculative decoding on moonshotai/Kimi-K3.
Serve with TokenSpeed (--speculative-algorithm DFLASH).
Provenance
| field | value |
|---|---|
| training run | kimi-k3-dflash2-la8-full-v2-nofcnorm |
| checkpoint | iter_0008001 (step 8000) — also best_checkpointed_iteration |
| TorchSpec source commit | e6c34e321d3885c7abafecb026a0295e762d3e12 |
| cluster | NVIDIA TSD GB300 Slurm, 4 inference nodes (2x TP8) + 2 training nodes |
fc_norm |
false (intended; the earlier e95a72e run had it unintentionally on) |
Training-side eval at this checkpoint
| metric | value |
|---|---|
eval/simulated_acc_len |
1.7784578371886477 |
eval/avg_acc |
0.4962460398674011 |
eval/avg_loss |
2.306995391845703 |
Architecture
DFlash2DraftModel, 6 layers, hidden 7168, bf16, 103 tensorsblock_size: 8(lookahead-8),attention_mode: mlatarget_layer_ids: [19, 37, 54, 66, 78, 90](6 aux hidden states)layer_types: 5xsliding_attention(window 4096) + 1xfull_attention- yarn RoPE,
rope_theta: 50000.0,factor: 32.0
Note: the validated serving path retains the full draft MLA KV cache in the unified full-attention storage group. Per-layer sliding-window metadata exists but the validated MLA backend does not consume it, so runs to date prove full-context MLA compute, not exact
sliding_attentioncompute.
Export file hashes (sha256)
config.json 3d846e0b26714bf662b5d816cc27f52ec3d6b8e0d93f087f2dcc3296b42284cf
model.safetensors fb54da2791a133473494a53ac8b5d93771a7a55de788ef924b3861179fc28a1f
- Downloads last month
- -
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support