DeepSeek-V4-Flash-oQ2.5e-mtp

This model was quantized using oQ (oMLX v0.5.0.dev1) mixed-precision quantization.

Quantization details

  • Model type: deepseek_v4
  • Bits: 2
  • Group size: 64
  • Format: MLX safetensors

Accuracy Benchmark

Runs used the oMLX standard accuracy benchmark harness with greedy decoding (temperature=0, top_p=1), thinking disabled, and batch size 32.

Type Model Size Avg MMLU 1000 Winogrande 1267 HumanEval 164 MBPP 300 MMLU-Pro 1000 KMMLU 500
Original precision DeepSeek V4 Flash (MLX engine) 148.7 GiB 78.09% 821/1000 82.10% 926/1267 73.09% 148/164 90.24% 249/300 83.00% 677/1000 67.70% 362/500 72.40%
MLX oQe DeepSeek-V4-Flash-oQ2.5e 103.1 GiB 74.56% 784/1000 78.40% 867/1267 68.43% 146/164 89.02% 234/300 78.00% 665/1000 66.50% 335/500 67.00%
GGUF Q4K 153.3 GiB 78.64% 820/1000 82.00% 946/1267 74.66% 147/164 89.63% 250/300 83.33% 682/1000 68.20% 370/500 74.00%
GGUF Q4K+IQ2XXS 90.9 GiB 69.85% 596/1000 59.60% 1024/1267 80.82% 146/164 89.02% 221/300 73.67% 624/1000 62.40% 268/500 53.60%
Downloads last month
1,860
Safetensors
Model size
33B params
Tensor type
BF16
·
U32
·
F32
·
U8
·
I32
·
MLX
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support