CloudSurf-4B-FC — official GGUF quants

Official GGUF quantizations of CloudSurf-4B-FC, a Gemma-4 E4B (effective-4B active; ~8.0B total params) function-calling / tool-use model by CloudSurf Software, quantized by us from the BF16 merged champion checkpoint. All benchmark scores, training details, and honesty disclosures live in the base model card — the scores there were measured on the BF16 weights, not on these quants. A quantized-model eval is pending and will be added here when it lands; until then, assume some quality loss below Q8_0, growing as the bit-width shrinks.

Every model file in this repo passed a 6-case tool-call smoke suite (simple call, multi-argument, tool selection, role:tool round-trip, irrelevance refusal, parallel calls) against llama-server --jinja before upload. A file that fails any case does not ship.

Files

File Size Use case Tool-call smoke
CloudSurf-4B-FC.Q4_K_M.gguf 5.3 GB Recommended default — fits 8 GB RAM/VRAM 6/6 PASS (2026-08-19)
CloudSurf-4B-FC.Q5_K_M.gguf 5.8 GB Balanced quality/size 6/6 PASS (2026-08-19)
CloudSurf-4B-FC.Q6_K.gguf 6.2 GB Quality pick for 8GB VRAM 6/6 PASS (2026-08-19)
CloudSurf-4B-FC.Q8_0.gguf 8.0 GB Near-lossless 6/6 PASS (2026-08-19)
CloudSurf-4B-FC.Q3_K_M.gguf 4.9 GB Low-RAM entry point 6/6 PASS (2026-08-19)
CloudSurf-4B-FC.BF16.gguf 15.1 GB Reference / requantize source 6/6 PASS (2026-08-19)

Vision (mmproj)

Pair any quant with a projector file for image input: mmproj-CloudSurf-4B-FC.F16.gguf (1.0 GB) or mmproj-CloudSurf-4B-FC.Q8_0.gguf (0.6 GB):

llama-server -m CloudSurf-4B-FC.Q4_K_M.gguf --mmproj mmproj-CloudSurf-4B-FC.F16.gguf --jinja -ngl 99

Function calling is this model's primary capability; vision rides along from the Gemma-4 base.

For imatrix (weighted) quants below Q3, see the community repo mradermacher/CloudSurf-4B-FC-i1-GGUF — an independent quantization of the same weights, credit to mradermacher for the fast pickup.

SHA-256 checksums for every file are in quant-manifest.json.

Running

llama.cpp (server, OpenAI-compatible tool calls)

--jinja is required — it enables the embedded chat template's tool-call rendering. Without it the server falls back to a generic template and tool calls will not parse.

llama-server -m CloudSurf-4B-FC.Q4_K_M.gguf --jinja -ngl 99 -c 16384 \
  --host 127.0.0.1 --port 8080 --alias cloudsurf-4b-fc
# then point any OpenAI client at http://127.0.0.1:8080/v1
# (chat/completions with tools in, tool_calls out)

llama.cpp (CLI)

llama-cli -m CloudSurf-4B-FC.Q4_K_M.gguf --jinja -ngl 99 -c 16384 \
  -p "You are a helpful assistant." -cnv

LM Studio

Search for CloudSurf-4B-FC-GGUF in the in-app model browser and download a quant. Enable the model's own prompt template (default) and use the OpenAI-compatible local server for tool calls.

Ollama

cat > Modelfile <<'EOF'
FROM ./CloudSurf-4B-FC.Q4_K_M.gguf
EOF
ollama create cloudsurf-4b-fc -f Modelfile
ollama run cloudsurf-4b-fc

The GGUF's embedded chat template carries the tool-call format; recent Ollama versions read it directly.

Notes and known quirks

  • Scores: the base card's BFCL V4 numbers (3-seed mean 55.73 as-registered / 54.91 matched-variant — both frames disclosed there) are BF16 measurements. No number in that card applies to these quants until the pending quantized-model eval is published here.
  • Params: "4B" is effective active parameters (the base's E4B convention); total is ~8.0B, which is what sizes these files.
  • </s> metadata quirk: GGUF conversions of this model can carry a stray </s> end-token metadata entry; llama.cpp detects and auto-corrects it at load (benign — observed and verified 2026-08-18).
  • Context: trained/evaluated at 131K-class context; -c 16384 above is a practical default, raise as your memory allows.

License

Apache-2.0, same as the base model. Gemma-4 base weights are subject to Google's Gemma terms as described in the base model card.

Downloads last month
266
GGUF
Model size
8B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cloudsurf-software/CloudSurf-4B-FC-GGUF

Quantized
(3)
this model