Clef-Flash GGUF, decision head kept at Q8_0, measured against BF16

GGUF quantizations of Cloudflare/clef-flash (9B decision model, Apache-2.0), made to run on a 16 GB machine. Text only: no mmproj here.

Clef-Flash does not generate text. It reads a state plus typed questions (choice, noul, score) and returns a probability for every allowed answer. llama-server serves it at POST /v1/systemone.

Resumo em português: quantizações GGUF do Clef-Flash da Cloudflare, com a cabeça de decisão preservada em Q8_0 e cada nível medido contra o BF16, inclusive num conjunto em português. As tabelas abaixo trazem a concordância com o BF16 e a memória que o servidor ocupa conforme o tamanho do pedido.

Files

File Size Same answer as BF16 Mean probability shift sha256
clef-flash-Q6_K.gguf 7.49 GB 98.6% 0.012 d84ab317f06b911fd6ee937625ae85e06827c565e27b37b4d524ca7456828629
clef-flash-Q5_K_M.gguf 6.59 GB 95.7% 0.031 81b6a65d508971e68dc0a7b35e27d852da172091e15bd221c959c358f5d10d5a
clef-flash-Q4_K_M.gguf 5.74 GB 94.4% 0.045 45e4565faef39c4347fbcb625896da6204f6f3bb37a1886de48f9a9b74cfbe69
clef-flash-Q3_K_M.gguf 4.75 GB 90.2% 0.077 4ab3b9ab663884bbfb3dd6e448de2c456bcf29ce0348a1e9c77990ce93ba7630

"Same answer" is the share of 2,630 decisions where the quantized file picks the same option as the BF16 GGUF. "Mean probability shift" takes, for each decision, the largest absolute difference in probability across its options, and averages that over the decisions.

How these were made

  • Source: the original safetensors from Cloudflare/clef-flash.
  • llama.cpp 0.6.0-dev, build 11459 (commit f498f864f), image ghcr.io/ggml-org/llama.cpp:full-cuda.
  • convert_hf_to_gguf.py --outtype bf16, then llama-quantize --tensor-type '^dec=q8_0' <bf16> <out> <LEVEL>.
  • The decision head (dec.blk.* and decision.*, about 245 MB in BF16) stays at Q8_0 at every level. Only the backbone follows the level.
  • No imatrix: llama.cpp cannot compute one for a decision model.

Quality, measured

One RTX 3090, llama-server from the same build, 4 parallel slots. The quantized levels answered every other item of the set; the BF16 row is BF16 on that same half.

Level typed-decisions (agreement with gold) PT-BR bench (balanced accuracy, mean of 7 tasks)
BF16 (18.16 GB) 0.711 0.654
Q6_K 0.701 0.662
Q5_K_M 0.701 0.642
Q4_K_M 0.704 0.639
Q3_K_M 0.678 0.651
  • LocalLLaMA/typed-decisions, test split: 200 cases, 1,000 decisions, English. The gold label is the average of three samples from a teacher model, so this is agreement with that teacher, not accuracy.
  • felhen-ai/ptbr-typed-decisions-bench, test split: 7 Portuguese tasks, up to 250 items per task here (1,630 in total), source labels, text cut at 3,500 characters, option order shuffled per item.
  • With 1,000 decisions the standard error is about 1.4 points, so Q4_K_M, Q5_K_M and Q6_K are not separable from BF16 on these two sets. The agreement column in the first table is the sharper measure. Q3_K_M changes one answer in ten.

Memory: it grows with the request

Clef evaluates the whole request (state plus questions) in one batch, so memory depends on how many tokens you send, not only on the file. Peak resident memory of llama-server on CPU (x86, -ngl 0, --parallel 1, -c = -b = -ub), one fresh server per row:

Level Context (-c) Request tokens Peak RSS
Q3_K_M 4096 374 5.7 GiB
Q3_K_M 2048 1,870 6.9 GiB
Q3_K_M 4096 3,400 8.5 GiB
Q4_K_M 4096 374 6.6 GiB
Q4_K_M 2048 1,870 7.8 GiB
Q4_K_M 4096 3,400 9.4 GiB
Q4_K_M 8192 6,800 13.0 GiB
Q4_K_M 16384 11,560 18.6 GiB
Q5_K_M 4096 374 7.4 GiB
Q5_K_M 4096 3,400 10.2 GiB
Q5_K_M 8192 6,800 13.8 GiB
Q5_K_M 16384 11,560 19.4 GiB

Rule of thumb from these rows: file size, plus about 0.13 GiB per 1,000 tokens of reserved context, plus about 0.93 GiB per 1,000 tokens of the request. Keep -c close to your largest request. Not measured: memory on Metal or CUDA at these request sizes, and Q6_K.

Run

llama-server -m clef-flash-Q4_K_M.gguf -c 4096 -b 4096 -ub 4096 --parallel 1

-c, -b and -ub go together, at least as large as your biggest request in tokens.

curl http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": "Mensagem do cliente: fui cobrado duas vezes pelo pedido da semana passada e ninguém respondeu.",
  "questions": {
    "equipe": {"type": "choice", "instructions": "Qual equipe deve atender?",
               "criteria": {"financeiro": null, "entrega": null, "tecnico": null}},
    "irritado": {"type": "noul", "instructions": "O cliente está irritado?"},
    "urgencia": {"type": "score", "instructions": "Qual a urgência?",
                 "criteria": ["pode esperar", "esta semana", "hoje", "agora"]}
  }
}'

Backends

  • CUDA and CPU (x86): the same request gets the same answer on both (checked with Q5_K_M).
  • Apple Silicon (Metal): use llama.cpp b11475 or newer. Older builds have a Metal bug that flattens Clef probabilities (#30064, fixed by #30100). We saw it with b11459 on an M5: the same file that answers 0.92 on CUDA answered 0.33 on Metal. On an older build, the fix's own test table shows GGML_METAL_FUSION_DISABLE=1 giving the right answer. We did not re-run on Metal after the fix.

License and credit

Apache-2.0, same as the original. Model by Cloudflare; this repo only converts and quantizes it and is not affiliated with Cloudflare.

Downloads last month
-
GGUF
Model size
9B params
Architecture
clef
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for felipeTromso/clef-flash-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(42)
this model