DeepSeek-V4-Flash (DSpark) โ IQ2_XXS + native MTP draft (GGUF)
Rhesadox GGUF of DeepSeek-V4-Flash with the DSpark speculative draft, re-quantized with the rhesadox deepseek-convert tool:
- Experts: IQ2_XXS (2-bit, imatrix-weighted) ยท down-projection: Q2_K ยท attention/projection/shared-experts/output: Q8_0
- MTP draft: native FP4 (DSpark, 3 layers, 256 experts/layer)
โ ๏ธ Split into 3 parts (HF's 50 GB per-file limit). Rejoin before use:
cat DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-nativeMTP-chat-v2-imatrix.gguf.part00 \
DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-nativeMTP-chat-v2-imatrix.gguf.part01 \
DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-nativeMTP-chat-v2-imatrix.gguf.part02 \
> model.gguf
Validated: main matches the official DeepSeek-V4-Flash API token oracle (golden 3/3 PASS); draft loads and passes SparkDrafter.init (diagnoseInit OK, 3 layers ร 256 experts).
Model tree for mrtib/DeepSeek-V4-Flash-DSpark-IQ2XXS
Base model
deepseek-ai/DeepSeek-V4-Flash