Qwen3.5-4B-MXFP4-GGUF

MXFP4 dense quant of unsloth/Qwen3.5-4B-GGUF, from llama.cpp PR timlikesai/llama.cpp#14 ("ggml: mxfp4 quantization, KV cache and Blackwell MMA"). New dense ftype LLAMA_FTYPE_MOSTLY_MXFP4 (=42).

Quantized with the UOS e8m0 block scale (MXAttention paper, arXiv:2607.24377): Q_max = 7.25 for E2M1, all 2D tensors MXFP4, token embeddings and output tensor at Q8_0 until MXFP8 lands.

Comparisons measured vs: Q4_0, Q4_1, Q4_K_M, Q8_0, BF16, UD-Q4_K_XL. Full data in the PR.

Downloads last month
289
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for timtimtimtimtim/Qwen3.5-4B-MXFP4-GGUF

Finetuned
Qwen/Qwen3.5-4B
Quantized
(10)
this model

Paper for timtimtimtimtim/Qwen3.5-4B-MXFP4-GGUF