DeepSeek-V4-Flash (DSpark) โ€” IQ2_XXS + native MTP draft (GGUF)

Rhesadox GGUF of DeepSeek-V4-Flash with the DSpark speculative draft, re-quantized with the rhesadox deepseek-convert tool:

  • Experts: IQ2_XXS (2-bit, imatrix-weighted) ยท down-projection: Q2_K ยท attention/projection/shared-experts/output: Q8_0
  • MTP draft: native FP4 (DSpark, 3 layers, 256 experts/layer)

โš ๏ธ Split into 3 parts (HF's 50 GB per-file limit). Rejoin before use:

cat DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-nativeMTP-chat-v2-imatrix.gguf.part00 \
    DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-nativeMTP-chat-v2-imatrix.gguf.part01 \
    DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-nativeMTP-chat-v2-imatrix.gguf.part02 \
    > model.gguf

Validated: main matches the official DeepSeek-V4-Flash API token oracle (golden 3/3 PASS); draft loads and passes SparkDrafter.init (diagnoseInit OK, 3 layers ร— 256 experts).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for mrtib/DeepSeek-V4-Flash-DSpark-IQ2XXS

Finetuned
(19)
this model