DeepSeek V4 Flash INT4 Dense (8GB GPU Mode)

This repository contains the 4-bit quantized dense weights (attention_dense_layers_q4.bin ~3.9 GB) for DeepSeek V4 Flash, optimized to run in 8 GB VRAM GPUs on the Moecher Inference Engine.

Running with Moecher (8GB GPUs)

moecher.exe --manifest models/deepseek_v4_flash_q4/moecher_manifest.json --max-vram 6 --dram-cache-gb 64 --quiet
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support