Instructions to use felipeTromso/clef-flash-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use felipeTromso/clef-flash-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf felipeTromso/clef-flash-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf felipeTromso/clef-flash-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf felipeTromso/clef-flash-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf felipeTromso/clef-flash-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf felipeTromso/clef-flash-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf felipeTromso/clef-flash-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf felipeTromso/clef-flash-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf felipeTromso/clef-flash-GGUF:Q4_K_M
Use Docker
docker model run hf.co/felipeTromso/clef-flash-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use felipeTromso/clef-flash-GGUF with Ollama:
ollama run hf.co/felipeTromso/clef-flash-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use felipeTromso/clef-flash-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf felipeTromso/clef-flash-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "felipeTromso/clef-flash-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use felipeTromso/clef-flash-GGUF with Docker Model Runner:
docker model run hf.co/felipeTromso/clef-flash-GGUF:Q4_K_M
- Lemonade
How to use felipeTromso/clef-flash-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull felipeTromso/clef-flash-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.clef-flash-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use felipeTromso/clef-flash-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf felipeTromso/clef-flash-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default felipeTromso/clef-flash-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use felipeTromso/clef-flash-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf felipeTromso/clef-flash-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "felipeTromso/clef-flash-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Clef-Flash GGUF, decision head kept at Q8_0, measured against BF16
GGUF quantizations of Cloudflare/clef-flash
(9B decision model, Apache-2.0), made to run on a 16 GB machine. Text only: no mmproj here.
Clef-Flash does not generate text. It reads a state plus typed questions (choice, noul,
score) and returns a probability for every allowed answer. llama-server serves it at
POST /v1/systemone.
Resumo em português: quantizações GGUF do Clef-Flash da Cloudflare, com a cabeça de decisão preservada em Q8_0 e cada nível medido contra o BF16, inclusive num conjunto em português. As tabelas abaixo trazem a concordância com o BF16 e a memória que o servidor ocupa conforme o tamanho do pedido.
Files
| File | Size | Same answer as BF16 | Mean probability shift | sha256 |
|---|---|---|---|---|
clef-flash-Q6_K.gguf |
7.49 GB | 98.6% | 0.012 | d84ab317f06b911fd6ee937625ae85e06827c565e27b37b4d524ca7456828629 |
clef-flash-Q5_K_M.gguf |
6.59 GB | 95.7% | 0.031 | 81b6a65d508971e68dc0a7b35e27d852da172091e15bd221c959c358f5d10d5a |
clef-flash-Q4_K_M.gguf |
5.74 GB | 94.4% | 0.045 | 45e4565faef39c4347fbcb625896da6204f6f3bb37a1886de48f9a9b74cfbe69 |
clef-flash-Q3_K_M.gguf |
4.75 GB | 90.2% | 0.077 | 4ab3b9ab663884bbfb3dd6e448de2c456bcf29ce0348a1e9c77990ce93ba7630 |
"Same answer" is the share of 2,630 decisions where the quantized file picks the same option as the BF16 GGUF. "Mean probability shift" takes, for each decision, the largest absolute difference in probability across its options, and averages that over the decisions.
How these were made
- Source: the original safetensors from
Cloudflare/clef-flash. - llama.cpp
0.6.0-dev, build 11459 (commitf498f864f), imageghcr.io/ggml-org/llama.cpp:full-cuda. convert_hf_to_gguf.py --outtype bf16, thenllama-quantize --tensor-type '^dec=q8_0' <bf16> <out> <LEVEL>.- The decision head (
dec.blk.*anddecision.*, about 245 MB in BF16) stays at Q8_0 at every level. Only the backbone follows the level. - No imatrix: llama.cpp cannot compute one for a decision model.
Quality, measured
One RTX 3090, llama-server from the same build, 4 parallel slots. The quantized levels
answered every other item of the set; the BF16 row is BF16 on that same half.
| Level | typed-decisions (agreement with gold) | PT-BR bench (balanced accuracy, mean of 7 tasks) |
|---|---|---|
| BF16 (18.16 GB) | 0.711 | 0.654 |
| Q6_K | 0.701 | 0.662 |
| Q5_K_M | 0.701 | 0.642 |
| Q4_K_M | 0.704 | 0.639 |
| Q3_K_M | 0.678 | 0.651 |
- LocalLLaMA/typed-decisions, test split: 200 cases, 1,000 decisions, English. The gold label is the average of three samples from a teacher model, so this is agreement with that teacher, not accuracy.
- felhen-ai/ptbr-typed-decisions-bench, test split: 7 Portuguese tasks, up to 250 items per task here (1,630 in total), source labels, text cut at 3,500 characters, option order shuffled per item.
- With 1,000 decisions the standard error is about 1.4 points, so Q4_K_M, Q5_K_M and Q6_K are not separable from BF16 on these two sets. The agreement column in the first table is the sharper measure. Q3_K_M changes one answer in ten.
Memory: it grows with the request
Clef evaluates the whole request (state plus questions) in one batch, so memory depends on
how many tokens you send, not only on the file. Peak resident memory of llama-server on CPU
(x86, -ngl 0, --parallel 1, -c = -b = -ub), one fresh server per row:
| Level | Context (-c) |
Request tokens | Peak RSS |
|---|---|---|---|
| Q3_K_M | 4096 | 374 | 5.7 GiB |
| Q3_K_M | 2048 | 1,870 | 6.9 GiB |
| Q3_K_M | 4096 | 3,400 | 8.5 GiB |
| Q4_K_M | 4096 | 374 | 6.6 GiB |
| Q4_K_M | 2048 | 1,870 | 7.8 GiB |
| Q4_K_M | 4096 | 3,400 | 9.4 GiB |
| Q4_K_M | 8192 | 6,800 | 13.0 GiB |
| Q4_K_M | 16384 | 11,560 | 18.6 GiB |
| Q5_K_M | 4096 | 374 | 7.4 GiB |
| Q5_K_M | 4096 | 3,400 | 10.2 GiB |
| Q5_K_M | 8192 | 6,800 | 13.8 GiB |
| Q5_K_M | 16384 | 11,560 | 19.4 GiB |
Rule of thumb from these rows: file size, plus about 0.13 GiB per 1,000 tokens of reserved
context, plus about 0.93 GiB per 1,000 tokens of the request. Keep -c close to your largest
request. Not measured: memory on Metal or CUDA at these request sizes, and Q6_K.
Run
llama-server -m clef-flash-Q4_K_M.gguf -c 4096 -b 4096 -ub 4096 --parallel 1
-c, -b and -ub go together, at least as large as your biggest request in tokens.
curl http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' -d '{
"state": "Mensagem do cliente: fui cobrado duas vezes pelo pedido da semana passada e ninguém respondeu.",
"questions": {
"equipe": {"type": "choice", "instructions": "Qual equipe deve atender?",
"criteria": {"financeiro": null, "entrega": null, "tecnico": null}},
"irritado": {"type": "noul", "instructions": "O cliente está irritado?"},
"urgencia": {"type": "score", "instructions": "Qual a urgência?",
"criteria": ["pode esperar", "esta semana", "hoje", "agora"]}
}
}'
Backends
- CUDA and CPU (x86): the same request gets the same answer on both (checked with Q5_K_M).
- Apple Silicon (Metal): use llama.cpp b11475 or newer. Older builds have a Metal bug
that flattens Clef probabilities
(#30064, fixed by
#30100). We saw it with b11459 on an
M5: the same file that answers 0.92 on CUDA answered 0.33 on Metal. On an older build, the
fix's own test table shows
GGML_METAL_FUSION_DISABLE=1giving the right answer. We did not re-run on Metal after the fix.
License and credit
Apache-2.0, same as the original. Model by Cloudflare; this repo only converts and quantizes it and is not affiliated with Cloudflare.
- Downloads last month
- -
3-bit
4-bit
5-bit
6-bit