Instructions to use Jundot/DeepSeek-V4-Flash-oQ2.5e-mtp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Jundot/DeepSeek-V4-Flash-oQ2.5e-mtp with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir DeepSeek-V4-Flash-oQ2.5e-mtp Jundot/DeepSeek-V4-Flash-oQ2.5e-mtp
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
DeepSeek-V4-Flash-oQ2.5e-mtp
This model was quantized using oQ (oMLX v0.5.0.dev1) mixed-precision quantization.
Quantization details
- Model type: deepseek_v4
- Bits: 2
- Group size: 64
- Format: MLX safetensors
Accuracy Benchmark
Runs used the oMLX standard accuracy benchmark harness with greedy decoding (temperature=0, top_p=1), thinking disabled, and batch size 32.
| Type | Model | Size | Avg | MMLU 1000 | Winogrande 1267 | HumanEval 164 | MBPP 300 | MMLU-Pro 1000 | KMMLU 500 |
|---|---|---|---|---|---|---|---|---|---|
| Original precision | DeepSeek V4 Flash (MLX engine) | 148.7 GiB | 78.09% | 821/1000 82.10% | 926/1267 73.09% | 148/164 90.24% | 249/300 83.00% | 677/1000 67.70% | 362/500 72.40% |
| MLX oQe | DeepSeek-V4-Flash-oQ2.5e | 103.1 GiB | 74.56% | 784/1000 78.40% | 867/1267 68.43% | 146/164 89.02% | 234/300 78.00% | 665/1000 66.50% | 335/500 67.00% |
| GGUF | Q4K | 153.3 GiB | 78.64% | 820/1000 82.00% | 946/1267 74.66% | 147/164 89.63% | 250/300 83.33% | 682/1000 68.20% | 370/500 74.00% |
| GGUF | Q4K+IQ2XXS | 90.9 GiB | 69.85% | 596/1000 59.60% | 1024/1267 80.82% | 146/164 89.02% | 221/300 73.67% | 624/1000 62.40% | 268/500 53.60% |
- Downloads last month
- 1,860
Model size
33B params
Tensor type
BF16
·
U32 ·
F32 ·
U8 ·
I32 ·
Hardware compatibility
Log In to add your hardware
2-bit
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support