Llamacpp Quantizations of DeepSeek-V4-Flash-0731 by deepseek-ai

Using llama.cpp release b10173 for quantization.

Original model: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731

This model is in MXFP4 and as such has only been provided in MXFP4 format!

Other sizes may be provided after some investigation.

Model details:

  • Parameter count: 284B
  • Input support: text
  • MTP: no
  • imatrix: no (for now)

How to run

Prompt format

No prompt format found

Download the MXFP4 files:

Filename Quant type File Size Split Description
DeepSeek-V4-Flash-0731-MXFP4.gguf MXFP4 156.38GB true Original quality.

Downloading using the Hugging Face CLI

Click to view download instructions

First, make sure you have the Hugging Face CLI installed:

pip install -U "huggingface_hub[cli]"

The files marked true in the Split column above are stored as multiple parts in a folder. To download all the parts to a local folder, run:

hf download bartowski/DeepSeek-V4-Flash-0731-GGUF --include "DeepSeek-V4-Flash-0731-MXFP4/*" --local-dir ./

You can either specify a new local-dir (DeepSeek-V4-Flash-0731-MXFP4) or download them all in place (./)

How to run

These quants run with llama.cpp - installable in one line via llama.app:

curl -LsSf https://llama.app/install.sh | sh
llama-server -hf bartowski/DeepSeek-V4-Flash-0731:MXFP4

llama-server includes a built-in chat web UI, served at http://localhost:8080 by default.

These quants were made with llama.cpp release b10173 - if this model's architecture is newly supported, you'll need that release or newer to run them.

They also work in: LM Studio · koboldcpp · ramalama · Jan AI · Text Generation Web UI · LoLLMs · Atomic Chat

Credits

Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset.

Thank you ZeroWw for the inspiration to experiment with embed/output.

Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

Downloads last month
-
GGUF
Model size
284B params
Architecture
deepseek4
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bartowski/DeepSeek-V4-Flash-0731-GGUF

Quantized
(43)
this model