DSv4-Flash aligned MTP (non-expert sidecar)

On-policy chain-alignment of mtp.0's non-expert parameters (74.3M, bf16) for DeepSeek-V4-Flash-0731, trained teacher-forced on the model's own greedy generations to improve chained (depth-k) draft acceptance.

Measured (single M3 Ultra, 4-bit pack, fixed depth-3 chain):

  • conditional acceptance d2 43.2% -> 51.1%, d3 9.8% -> 22.2%
  • tokens/verify-cycle 2.25 -> 2.31

Load after the 4-bit pack by dequantizing mtp.0's non-expert linears to bf16 and applying these tensors (see train_align.py / serving sidecar in github.com/avlp12/dsv4flash_tp2_stack). Speculative decoding is lossless, so this changes speed only, never outputs.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for avlp12/dsv4flash-mtp-aligned

Finetuned
(26)
this model