DSv4-Flash aligned MTP (non-expert sidecar)
On-policy chain-alignment of mtp.0's non-expert parameters (74.3M, bf16) for
DeepSeek-V4-Flash-0731, trained teacher-forced on the model's own greedy
generations to improve chained (depth-k) draft acceptance.
Measured (single M3 Ultra, 4-bit pack, fixed depth-3 chain):
- conditional acceptance d2 43.2% -> 51.1%, d3 9.8% -> 22.2%
- tokens/verify-cycle 2.25 -> 2.31
Load after the 4-bit pack by dequantizing mtp.0's non-expert linears to bf16
and applying these tensors (see train_align.py / serving sidecar in
github.com/avlp12/dsv4flash_tp2_stack). Speculative decoding is lossless, so
this changes speed only, never outputs.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for avlp12/dsv4flash-mtp-aligned
Base model
deepseek-ai/DeepSeek-V4-Flash