DeepSeek-V4-Flash DSpark draft-only FP8

This repository contains only the mtp.* DSpark draft tensors from deepseek-ai/DeepSeek-V4-Flash-DSpark. Its routed-expert E2M1 FP4 tensors were losslessly represented as block-scaled E4M3 FP8 tensors using SGLang's cast_e2m1fn_to_e4m3fn conversion.

Use it with an FP8 target checkpoint:

python3 -m sglang.launch_server \
  --model-path sgl-project/DeepSeek-V4-Flash-FP8 \
  --speculative-algorithm DSPARK \
  --speculative-draft-model-path AtlasCloud/DeepSeek-V4-Flash-DSpark-FP8-dspark_only

The target checkpoint must provide the compatible tokenizer, embedding, and LM head; this draft-only checkpoint intentionally excludes them.

Downloads last month
1,076
Safetensors
Model size
20B params
Tensor type
F32
BF16
F8_E8M0
F8_E4M3
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for AtlasCloud/DeepSeek-V4-Flash-DSpark-FP8-dspark_only

Quantized
(11)
this model