Clef-Flash GGUFs on bloomery: which quant to pick
This page has no weights. It shows how close each of bartowski's Clef-Flash GGUFs answers to the official BF16 model, when run with bloomery. Use it to pick a file for your GPU.
This is an unofficial page. It is not made by or affiliated with Cloudflare. The model is Cloudflare/clef-flash (Apache-2.0).
Why bloomery
Clef is a decision model. One prompt pass gives the backbone's hidden states, and a small "joint schema head" turns
them into a probability for every allowed option. A GGUF holds only the backbone. llama.cpp has no way to run the
head. bloomery runs both and serves the release's SystemOne API (POST /v1/systemone).
Agreement with the official model
We ran 15 requests (8 English, 7 Korean; 31 questions) through Cloudflare's own Python code at BF16 and through bloomery with each file. "Top equal" counts questions where both pick the same option. "Max |Δp|" is the largest probability gap on the chosen option.
| File | Size | Top equal | Korean | Max |Δp| | Prompt, tok/s |
|---|---|---|---|---|---|
| Q8_0 | 9.55 GB | 30/31 | 15/15 | 0.014 | 3,289 |
| Q6_K_L | 8.11 GB | 31/31 | 15/15 | 0.042 | 3,333 |
| Q6_K | 7.79 GB | 31/31 | 15/15 | 0.025 | 3,363 |
| Q6_K_S | 7.51 GB | 31/31 | 15/15 | 0.034 | 3,383 |
| Q5_K_M | 6.88 GB | 31/31 | 15/15 | 0.052 | 3,862 |
| Q5_K_S | 6.50 GB | 31/31 | 15/15 | 0.057 | 3,998 |
| Q4_K_L | 6.20 GB | 30/31 | 15/15 | 0.077 | 4,128 |
| Q4_K_M | 5.84 GB | 29/31 | 15/15 | 0.072 | 4,322 |
| Q4_K_S | 5.48 GB | 29/31 | 15/15 | 0.078 | 4,526 |
| Q3_K_L | 4.66 GB | 30/31 | 15/15 | 0.080 | 4,167 |
| Q3_K_M | 4.48 GB | 30/31 | 15/15 | 0.102 | 4,087 |
| Q3_K_S | 4.26 GB | 29/31 | 15/15 | 0.136 | 4,033 |
- Every miss is a question where the official model is itself almost a tie (0.474 against 0.471) or near 0.5 (0.512). So the top choice is stable down to Q3_K_S. The probabilities drift as the file gets smaller.
- If you use the probabilities, not only the top choice, pick Q6_K or Q5_K_M.
- Prompt speed: one RTX A6000, a 4,067-token request, one request at a time. These are functional runs, not benchmark runs; the same file moves about 3–4 % between runs. The head runs on the CPU: a median of 22 ms per request, about 100 ms at 4,000 tokens.
- Not supported yet: the IQ files, Q2_K, Q4_0, Q4_1 and BF16. bloomery refuses them by name at load.
The full table, the requests and the scripts are in the bloomery repository:
tools/ref/clef/agreement.md.
Run it
You need a Linux x86-64 host with an NVIDIA RTX 30-series or A-series GPU (sm_86) and the bloomery toolchain (build guide).
hf download bartowski/Cloudflare_clef-flash-GGUF Cloudflare_clef-flash-Q5_K_M.gguf --local-dir ~/models/clef-flash
hf download Cloudflare/clef-flash joint_head.safetensors joint_head_config.json --local-dir ~/models/clef-flash/hf
cargo oxide build --arch sm_86 -- -p bloomery-gpu-gates --features clef --release --bin bloomery_serve_clef
target/release/bloomery_serve_clef --model ~/models/clef-flash/Cloudflare_clef-flash-Q5_K_M.gguf \
--head ~/models/clef-flash/hf/joint_head.safetensors --port 8091
curl -s http://127.0.0.1:8091/v1/systemone -d '{"model": "clef-flash",
"state": "User: what is the weather in Seoul tomorrow? Tools available: web_search, calculator, calendar.",
"questions": {"tool": {"type": "choice", "instructions": "Which tool should the agent call next?",
"criteria": {"web_search": "Look something up online", "calculator": "Do arithmetic",
"calendar": "Read or write events", "none": "Answer directly"}}}}'
The request and response follow the release's SystemOne format. Each response also has timings (prompt_n,
prompt_ms, head_ms).
Limits
- Text states only (no images or video).
- One request at a time.
- sm_86 GPUs only (RTX 3090, RTX A6000).
Credits
Model: Cloudflare (Clef-Flash, Apache-2.0). GGUFs: bartowski. Engine: bloomery (MIT). AI assistants helped write bloomery and this page; every number above was measured.