Qwen3.5-2B-Agent GGUF

GGUF conversion of Ma7ee7/Qwen3.5-2B-Agent for llama.cpp-compatible applications.

The multimodal projector is included as mmproj-Qwen3.5-2B-Agent-Q8_0.gguf.

Converted with llama.cpp commit 720d7fa4097f76e5d0eade5a92c1df87c1faf9d9.

Files

File Size Description
Qwen3.5-2B-Agent-BF16.gguf 3.63 GiB High-fidelity source GGUF for archival use or requantization.
Qwen3.5-2B-Agent-Q8_0.gguf 1.93 GiB Near-original quality with moderate compression.
Qwen3.5-2B-Agent-Q6_K.gguf 1.50 GiB Very high quality with lower memory use than Q8_0.
Qwen3.5-2B-Agent-Q5_K_M.gguf 1.35 GiB Strong quality-to-size balance.
Qwen3.5-2B-Agent-Q4_K_M.gguf 1.22 GiB Recommended default for most local inference.
mmproj-Qwen3.5-2B-Agent-Q8_0.gguf 347.77 MiB Multimodal projector required for image input.

Recommendation

Start with Q4_K_M. Use Q5_K_M or Q6_K when you have more memory and want higher fidelity. Q8_0 is the largest quantized option in this repository.

llama.cpp

llama-cli -hf Ma7ee7/Qwen3.5-2B-Agent-GGUF:Q4_K_M -cnv --jinja

For image input, download the matching mmproj-*.gguf file and pass it with --mmproj.

Downloads last month
321
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Ma7ee7/Qwen3.5-2B-Agent-GGUF

Finetuned
Qwen/Qwen3.5-2B
Quantized
(1)
this model