DeepSeek V4 Flash INT4 Dense (8GB GPU Mode)
This repository contains the 4-bit quantized dense weights (attention_dense_layers_q4.bin ~3.9 GB) for DeepSeek V4 Flash, optimized to run in 8 GB VRAM GPUs on the Moecher Inference Engine.
Running with Moecher (8GB GPUs)
moecher.exe --manifest models/deepseek_v4_flash_q4/moecher_manifest.json --max-vram 6 --dram-cache-gb 64 --quiet