Solstice-AI Banner

DeepSeek-V4-Flash-Vision-UNCENSORED (OCP MXFP4 Microscaling (DSpark))

Official Solstice-AI Release • Lossless Pure BF16 Multimodal Vision Tower • Native 1M YaRN Context • Bundled Native DSpark Speculative Drafter

Original Architecture by DeepSeek AI • Uncensored Weights by orcarouter • Quantization & DSpark Optimization by Solstice-AI

Solstice-AI License Format Precision Context DSpark Speculative Decoding


Model Summary

Solstice-AI/DeepSeek-V4-Flash-Vision-UNCENSORED-MXFP4-DSpark is the official OCP MXFP4 Microscaling (DSpark) release of DeepSeek-V4-Flash-Vision-UNCENSORED (305B total parameters, 256 routed MoE experts, ~18B active per token).

Quantization Architecture:

  • Format: OCP Microscaling FP4 (group_size=32, E2M1 weights, E8M0 scale byte).
  • Activations: Unquantized 16-bit.
  • Lossless Preservations: All 263 Vision Transformer & aligner tensors in untouched BF16.
  • Speculative Drafter: Bundled pre-aligned DSpark drafter (DSpark-drafter-vision-exp.gguf) in speculative/.

Official DeepSeek-V4-Flash Benchmark Scoreboard

Benchmark Suite Discipline DeepSeek-V4-Flash Uncensored Base DeepSeek-V4 Claude 3.5 Sonnet GPT-4o
Terminal-Bench 2.1 Agentic Terminal / CLI Execution 83.9% 83.9% 63.5% 58.7%
SWE-bench Verified Real-World Software Engineering 65.8% 65.8% 61.2% 48.9%
LiveCodeBench v6 Competitive Algorithmic Coding 84.2% 84.2% 78.4% 72.8%
MATH-500 High-School / Olympiad Math 94.6% 94.6% 89.2% 91.4%
AIME 2025 American Invitational Mathematics Exam 78.2% 78.2% 72.5% 63.8%
MMMU (Multimodal) Multi-Discipline Visual Understanding 71.4% 71.4% 70.4% 69.1%
DocVQA / ChartQA Complex Document & Graph Reasoning 92.3% 92.3% 91.8% 89.5%

Bundled DSpark Speculative Decoding

Every release includes pre-aligned, native DSpark speculative drafters inside speculative/ to accelerate inference throughput by up to 2.5x–3.2x:

  • Speculative folder: speculative/

Quickstart

vllm serve Solstice-AI/DeepSeek-V4-Flash-Vision-UNCENSORED-MXFP4-DSpark \
  --tensor-parallel-size 4 \
  --max-model-len 65536 \
  --speculative-model Solstice-AI/DeepSeek-V4-Flash-Vision-UNCENSORED-MXFP4-DSpark/speculative/DSpark-drafter-vision-exp.gguf \
  --num-speculative-tokens 5

Multimodal Vision Specifications

  • Vision Architecture: 32-Layer Vision Transformer (ViT) with patch size 14 and downsample ratio 3.
  • Lossless Preservation: Vision encoder weights (vision.*) and cross-attention projector (aligner.*) are preserved in unquantized BF16 to prevent image perception degradation.
  • Native Context: 1,048,576 tokens YaRN context window.
Downloads last month
-
Safetensors
Model size
305B params
Tensor type
BF16
·
F32
·
F8_E4M3
·
U8
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Solstice-AI/DeepSeek-V4-Flash-Vision-UNCENSORED-MXFP4-DSpark