DASH-Q

phi-4 - DASH-Q 2-bit GGUF

llama.cpp GGUF files of microsoft/phi-4 quantized with DASH-Q, at several 2-bit-class sizes. Every file uses only standard llama.cpp tensor types (no tensor above 4 bits) and loads in any recent llama.cpp build.

File Type Size Bits / weight
phi-4-DASHQ-IQ2_XXS.gguf IQ2_XXS 4.66 GB 2.54
phi-4-DASHQ-IQ2_XS.gguf IQ2_XS 5.18 GB 2.83
phi-4-DASHQ-IQ2_M.gguf IQ2_M 5.61 GB 3.07
phi-4-DASHQ-Q2_K_XL.gguf Q2_K_XL 6.06 GB 3.31

Perplexity (lower is better)

Type Model Size WikiText-2 C4
F16 29.3 GB 5.78 11.27
IQ2_XXS llama.cpp IQ2_XXS (imatrix) 4.55 GB 7.59 13.94
IQ2_XXS DASH-Q IQ2_XXS 4.66 GB 7.38 13.80
IQ2_XS llama.cpp IQ2_XS (imatrix) 4.92 GB 7.01 13.05
IQ2_XS DASH-Q IQ2_XS 5.18 GB 6.56 12.47
IQ2_M llama.cpp IQ2_M (imatrix) 5.49 GB 6.46 12.28
IQ2_M DASH-Q IQ2_M 5.61 GB 6.27 11.97
Q2_K_XL llama.cpp Q2_K (imatrix) 5.92 GB 6.42 12.17
Q2_K_XL unsloth Q2_K_L 5.73 GB 6.74 12.61
Q2_K_XL DASH-Q Q2_K_XL 6.06 GB 6.18 11.87

llama-perplexity, context 2048; WikiText-2 test, C4 validation (256 x 2048 tokens).

Zero-shot accuracy (lm-eval-harness, 0-shot)

File arc_easy arc_challenge piqa hellaswag winogrande openbookqa commonsense_qa truthfulqa_mc2 lambada_openai avg
DASH-Q IQ2_XXS 77.5 53.1 79.8 75.0 75.1 43.8 72.7 54.2 75.8 67.44
DASH-Q IQ2_XS 76.3 56.1 81.0 78.4 76.4 46.2 76.0 55.9 74.7 69.01
DASH-Q IQ2_M 74.2 54.8 81.2 79.2 74.0 44.6 76.8 56.7 75.5 68.56
DASH-Q Q2_K_XL 74.4 55.5 81.4 80.2 76.9 44.6 76.2 56.4 74.2 68.86

Usage

llama-cli -m phi-4-DASHQ-IQ2_M.gguf -ngl 99 -c 8192

License

Inherits the license of the base model (microsoft/phi-4).

Downloads last month
347
GGUF
Model size
15B params
Architecture
phi3
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for jkim96/phi-4-DASHQ-Q2-GGUF

Base model

microsoft/phi-4
Quantized
(169)
this model