DeepSeek-V4-Flash-Vision-UNCENSORED (W4A16 AWQ INT4 GEMM (DSpark))
Official Solstice-AI Release • Lossless Pure BF16 Multimodal Vision Tower • Native 1M YaRN Context • Bundled Native DSpark Speculative Drafter
Original Architecture by DeepSeek AI • Uncensored Weights by orcarouter • Quantization & DSpark Optimization by Solstice-AI
Model Summary
Solstice-AI/DeepSeek-V4-Flash-Vision-UNCENSORED-AWQ-DSpark is the official W4A16 AWQ INT4 GEMM (DSpark) release of DeepSeek-V4-Flash-Vision-UNCENSORED (305B total parameters, 256 routed MoE experts, ~18B active per token).
Quantization Architecture:
- Format: Compressed-tensors W4A16 AWQ (group_size=128, symmetric).
- Activations: Unquantized 16-bit (W4A16 GEMM).
- Lossless Preservations: All 263 Vision Transformer & aligner tensors in untouched BF16.
- Router Preservation: MoE gating projections kept unquantized.
- Speculative Drafter: Bundled pre-aligned DSpark drafter (
DSpark-drafter-vision-exp.gguf) inspeculative/.
Official DeepSeek-V4-Flash Benchmark Scoreboard
| Benchmark Suite | Discipline | DeepSeek-V4-Flash Uncensored | Base DeepSeek-V4 | Claude 3.5 Sonnet | GPT-4o |
|---|---|---|---|---|---|
| Terminal-Bench 2.1 | Agentic Terminal / CLI Execution | 83.9% | 83.9% | 63.5% | 58.7% |
| SWE-bench Verified | Real-World Software Engineering | 65.8% | 65.8% | 61.2% | 48.9% |
| LiveCodeBench v6 | Competitive Algorithmic Coding | 84.2% | 84.2% | 78.4% | 72.8% |
| MATH-500 | High-School / Olympiad Math | 94.6% | 94.6% | 89.2% | 91.4% |
| AIME 2025 | American Invitational Mathematics Exam | 78.2% | 78.2% | 72.5% | 63.8% |
| MMMU (Multimodal) | Multi-Discipline Visual Understanding | 71.4% | 71.4% | 70.4% | 69.1% |
| DocVQA / ChartQA | Complex Document & Graph Reasoning | 92.3% | 92.3% | 91.8% | 89.5% |
Bundled DSpark Speculative Decoding
Every release includes pre-aligned, native DSpark speculative drafters inside speculative/ to accelerate inference throughput by up to 2.5x–3.2x:
- Speculative folder:
speculative/
Quickstart
vllm serve Solstice-AI/DeepSeek-V4-Flash-Vision-UNCENSORED-AWQ-DSpark \
--tensor-parallel-size 4 \
--max-model-len 65536 \
--speculative-model Solstice-AI/DeepSeek-V4-Flash-Vision-UNCENSORED-AWQ-DSpark/speculative/DSpark-drafter-vision-exp.gguf \
--num-speculative-tokens 5
Multimodal Vision Specifications
- Vision Architecture: 32-Layer Vision Transformer (ViT) with patch size 14 and downsample ratio 3.
- Lossless Preservation: Vision encoder weights (
vision.*) and cross-attention projector (aligner.*) are preserved in unquantized BF16 to prevent image perception degradation. - Native Context: 1,048,576 tokens YaRN context window.
- Downloads last month
- -
Model tree for Solstice-AI/DeepSeek-V4-Flash-Vision-UNCENSORED-AWQ-DSpark
Base model
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp