Clef-Flash GGUFs on bloomery: which quant to pick

This page has no weights. It shows how close each of bartowski's Clef-Flash GGUFs answers to the official BF16 model, when run with bloomery. Use it to pick a file for your GPU.

This is an unofficial page. It is not made by or affiliated with Cloudflare. The model is Cloudflare/clef-flash (Apache-2.0).

Why bloomery

Clef is a decision model. One prompt pass gives the backbone's hidden states, and a small "joint schema head" turns them into a probability for every allowed option. A GGUF holds only the backbone. llama.cpp has no way to run the head. bloomery runs both and serves the release's SystemOne API (POST /v1/systemone).

Agreement with the official model

We ran 15 requests (8 English, 7 Korean; 31 questions) through Cloudflare's own Python code at BF16 and through bloomery with each file. "Top equal" counts questions where both pick the same option. "Max |Δp|" is the largest probability gap on the chosen option.

File Size Top equal Korean Max |Δp| Prompt, tok/s
Q8_0 9.55 GB 30/31 15/15 0.014 3,289
Q6_K_L 8.11 GB 31/31 15/15 0.042 3,333
Q6_K 7.79 GB 31/31 15/15 0.025 3,363
Q6_K_S 7.51 GB 31/31 15/15 0.034 3,383
Q5_K_M 6.88 GB 31/31 15/15 0.052 3,862
Q5_K_S 6.50 GB 31/31 15/15 0.057 3,998
Q4_K_L 6.20 GB 30/31 15/15 0.077 4,128
Q4_K_M 5.84 GB 29/31 15/15 0.072 4,322
Q4_K_S 5.48 GB 29/31 15/15 0.078 4,526
Q3_K_L 4.66 GB 30/31 15/15 0.080 4,167
Q3_K_M 4.48 GB 30/31 15/15 0.102 4,087
Q3_K_S 4.26 GB 29/31 15/15 0.136 4,033
  • Every miss is a question where the official model is itself almost a tie (0.474 against 0.471) or near 0.5 (0.512). So the top choice is stable down to Q3_K_S. The probabilities drift as the file gets smaller.
  • If you use the probabilities, not only the top choice, pick Q6_K or Q5_K_M.
  • Prompt speed: one RTX A6000, a 4,067-token request, one request at a time. These are functional runs, not benchmark runs; the same file moves about 3–4 % between runs. The head runs on the CPU: a median of 22 ms per request, about 100 ms at 4,000 tokens.
  • Not supported yet: the IQ files, Q2_K, Q4_0, Q4_1 and BF16. bloomery refuses them by name at load.

The full table, the requests and the scripts are in the bloomery repository: tools/ref/clef/agreement.md.

Run it

You need a Linux x86-64 host with an NVIDIA RTX 30-series or A-series GPU (sm_86) and the bloomery toolchain (build guide).

hf download bartowski/Cloudflare_clef-flash-GGUF Cloudflare_clef-flash-Q5_K_M.gguf --local-dir ~/models/clef-flash
hf download Cloudflare/clef-flash joint_head.safetensors joint_head_config.json --local-dir ~/models/clef-flash/hf

cargo oxide build --arch sm_86 -- -p bloomery-gpu-gates --features clef --release --bin bloomery_serve_clef
target/release/bloomery_serve_clef --model ~/models/clef-flash/Cloudflare_clef-flash-Q5_K_M.gguf \
  --head ~/models/clef-flash/hf/joint_head.safetensors --port 8091

curl -s http://127.0.0.1:8091/v1/systemone -d '{"model": "clef-flash",
  "state": "User: what is the weather in Seoul tomorrow? Tools available: web_search, calculator, calendar.",
  "questions": {"tool": {"type": "choice", "instructions": "Which tool should the agent call next?",
    "criteria": {"web_search": "Look something up online", "calculator": "Do arithmetic",
                 "calendar": "Read or write events", "none": "Answer directly"}}}}'

The request and response follow the release's SystemOne format. Each response also has timings (prompt_n, prompt_ms, head_ms).

Limits

  • Text states only (no images or video).
  • One request at a time.
  • sm_86 GPUs only (RTX 3090, RTX A6000).

Credits

Model: Cloudflare (Clef-Flash, Apache-2.0). GGUFs: bartowski. Engine: bloomery (MIT). AI assistants helped write bloomery and this page; every number above was measured.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for midagedev/clef-flash-bloomery

Finetuned
Qwen/Qwen3.5-9B
Quantized
(22)
this model