Toolcall-2B — GGUF

Quantized builds of ajvikram/toolcall-2b, a 2B function-calling model fine-tuned from Qwen3.5-2B for local agent tool routing. Full results, training details and limitations are on the parent model's card.

File Size Use
toolcall-2b-Q4_K_M.gguf 1.22 GB Default. Smallest sensible quality loss, runs on a laptop CPU.
toolcall-2b-Q5_K_M.gguf 1.35 GB A little closer to full precision for modest extra memory.
toolcall-2b-Q8_0.gguf 1.93 GB Near-lossless; use when you have the memory.
toolcall-2b-f16.gguf 3.63 GB Unquantized source for making your own quants.

Measured on the benchmark harness (safetensors, bf16): 36.35 overall on BFCL v4 against 33.85 for the Qwen3.5-2B base, with every group ahead of the base. The quantized builds are not separately scored.

Run it

llama-server -m toolcall-2b-Q4_K_M.gguf --jinja -c 8192
ollama run hf.co/ajvikram/toolcall-2b-gguf:Q4_K_M

The model uses Qwen3.5's native XML tool-call format, so any client that already parses Qwen3.5 tool calls works unchanged:

<tool_call>
<function=get_weather>
<parameter=city>
Berlin
</parameter>
</function>
</tool_call>

Verified with llama-cli on CPU: the Q4_K_M build loads, generates at roughly 33 tokens per second on an ARM CPU, and returns the call above for a get_weather tool given "What is the weather in Berlin?".

Thinking is off by default, matching how the model was trained and evaluated.

Notes

  • Built with llama.cpp (September 2026), which added Qwen3.5 conversion support; older builds cannot convert this architecture.
  • These are text-only builds. The base architecture is vision-capable, but this model was trained and evaluated purely on text tool calling.
Downloads last month
-
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ajvikram/toolcall-2b-gguf

Finetuned
Qwen/Qwen3.5-2B
Quantized
(2)
this model