Instructions to use Jakevin/clef-flash-mixed-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Jakevin/clef-flash-mixed-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Jakevin/clef-flash-mixed-GGUF # Run inference directly in the terminal: llama cli -hf Jakevin/clef-flash-mixed-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Jakevin/clef-flash-mixed-GGUF # Run inference directly in the terminal: llama cli -hf Jakevin/clef-flash-mixed-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Jakevin/clef-flash-mixed-GGUF # Run inference directly in the terminal: ./llama-cli -hf Jakevin/clef-flash-mixed-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Jakevin/clef-flash-mixed-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf Jakevin/clef-flash-mixed-GGUF
Use Docker
docker model run hf.co/Jakevin/clef-flash-mixed-GGUF
- LM Studio
- Jan
- Ollama
How to use Jakevin/clef-flash-mixed-GGUF with Ollama:
ollama run hf.co/Jakevin/clef-flash-mixed-GGUF
- Unsloth Desktop
- Pi
How to use Jakevin/clef-flash-mixed-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Jakevin/clef-flash-mixed-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Jakevin/clef-flash-mixed-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Jakevin/clef-flash-mixed-GGUF with Docker Model Runner:
docker model run hf.co/Jakevin/clef-flash-mixed-GGUF
- Lemonade
How to use Jakevin/clef-flash-mixed-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Jakevin/clef-flash-mixed-GGUF
Run and chat with the model
lemonade run user.clef-flash-mixed-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use Jakevin/clef-flash-mixed-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Jakevin/clef-flash-mixed-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Jakevin/clef-flash-mixed-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Jakevin/clef-flash-mixed-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Jakevin/clef-flash-mixed-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Jakevin/clef-flash-mixed-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Clef-Flash mixed-precision GGUF (measured-KL, text only)
Mixed precision, not ternary. Tensors use IQ2_XXS, IQ2_S, Q2_K, IQ3_XXS, Q3_K, Q4_K and Q8_0; no ternary format (TQ1_0/TQ2_0) is used. This repo was renamed from Jakevin/clef-flash-ternary-GGUF on 2026-10-09; the old URL redirects here and the v1.0 tag is unchanged.
Version v1.0 (2026-10-09). Unofficial post-training quantization of
Cloudflare/clef-flash
revision 17f0b0ad64efb65d273590632833508766b2aae6. This is not an official
Cloudflare release and is not endorsed by Cloudflare or the Qwen team.
Licensed Apache-2.0 like the original (LICENSE). NOTICE.md lists the changes.
The file is the text backbone only. There is no vision tower. llama.cpp by
itself does not produce Clef decisions: the joint schema head runs in Python
on the backbone's final hidden states, and the option rows come from the
original bf16 lm_head. The GGUF output.weight is Q2_K and the head does
not read it.
general.architecture is qwen35. The llama.cpp clef graph does not export
t_h_nextn, which is the hidden state this recipe reads, so the file is the
Qwen3.5 text model rather than architecture clef. general.name is
Snap Text because the conversion directory had that name. The weights are
Clef-Flash text.
Size
| bytes | |
|---|---|
clef-flash-mkl-3.30GB.gguf |
3.300 GB (3,299,992,800) |
| original bf16 release | 19.06 GB (18.82 GB shards + 0.24 GB head) |
Scores
Author's private frozen development splits (kev). One seed. No confidence interval was computed.
| suite | n (clean) | bf16 acc | this model | retained |
|---|---|---|---|---|
| decision-v7 | 1264 | 0.8861 | 0.8813 | 99.5% |
| transfer-v9 | 1046 | 0.8011 | 0.7859 | 98.1% |
| suite | Brier | NLL | ECE |
|---|---|---|---|
| decision-v7 | 0.1755 | 0.3382 | 0.0278 |
| transfer-v9 | 0.3061 | 0.6042 | 0.0421 |
Retention is this model's clean accuracy divided by the bf16 clean accuracy on the same suite.
decision-v7 can be optimistic. The imatrix and the KL allocation both used decision-v7 records: 128 calibration records, seed 1234, for the imatrix, and the first 32 of that draw (10,383 tokens) for the per-tensor KL. transfer-v9 was not used for calibration.
These are not the public Decision Index or Typesafe numbers.
Method
Per-tensor measured-KL sensitivity, then a knapsack that assigns each measured
tensor a ggml type from IQ2_XXS, IQ2_S, Q2_K, IQ3_XXS, Q3_K, Q4_K, and Q8_0.
The shipped file is llama-quantize --tensor-type with that imatrix.
alloc/selection.json is the knapsack result. alloc/tensor-types.txt is the
pattern file passed to the quantizer (it pins ssm_alpha and ssm_beta to
bf16; the other assignments are exact tensor names).
Counts in the GGUF (427 tensors):
| type | tensors | what |
|---|---|---|
| IQ2_XXS | 68 | body linears |
| IQ2_S | 23 | body linears |
| IQ3_XXS | 35 | body linears and token_embd |
| Q2_K | 16 | 15 body linears plus output.weight |
| Q3_K | 13 | body linears |
| Q4_K | 36 | body linears |
| Q8_0 | 11 | body linears |
| BF16 | 48 | ssm_alpha, ssm_beta (24 each) |
| F32 | 177 | norms, ssm_conv1d, ssm_a, ssm_dt |
Norms, conv, ssm_alpha, and ssm_beta are not quantized. In the file, norms
and conv are F32; ssm_alpha and ssm_beta are BF16. output.weight is
fixed at Q2_K and is not read by the classification head. token_embd is
IQ3_XXS (IQ2_XXS asserts an imatrix, and llama-quantize does not pass one
for the embedding).
Sum of the per-tensor KL values used by the knapsack (not a model KL): 0.009015. Dry-run file estimate was 3.300 GB.
Other files of the same source model
Listed side by side. Different formats and different calibration. This card does not say which allocation is better.
| release | what it is | size | decision-v7 | transfer-v9 |
|---|---|---|---|---|
| this repo, v1.0 | GGUF, measured-KL mixed types | 3.300 GB | 0.8813 (99.5% of bf16) | 0.7859 (98.1% of bf16) |
| Jakevin/clef-flash-ternary-mlx v2.0 | MLX packed mixed-bit GPTQ | 3.13 GB | 0.8766 (98.9% of bf16) | 0.7361 (91.9% of bf16) |
bartowski/Cloudflare_clef-flash-GGUF Cloudflare_clef-flash-Q2_K.gguf |
GGUF Q2_K | 4.19 GB (4,194,031,776 bytes) | not scored here | not scored here |
The MLX v2.0 scores are the ones on that repo's model card (same bf16 references, 0.8861 and 0.8011). Its calibration and packed format are not this GGUF's. The bartowski Q2_K file size is the file on that repo. We did not evaluate it.
Use
The joint head needs torch, safetensors, numpy, and transformers.
hsdump needs llama.cpp b11407 or later with text-only qwen35. Build it
from hsdump.cpp in this folder. The two embedding calls are already in that
llama.cpp (src/llama-ext.h); they are not in the installed llama.h, so the
cpp file declares them. No llama.cpp patch is required at b11407 or later.
output.weight inside the GGUF is not the classification head. Pass
--base as a checkout of Cloudflare/clef-flash (the same revision as above).
Only lm_head.weight is read from it.
c++ -std=c++17 -O2 \
-I llama.cpp/include -I llama.cpp/ggml/include \
hsdump.cpp -o hsdump \
-L /path/to/llama.cpp/build/bin -lllama \
-Wl,-rpath,/path/to/llama.cpp/build/bin
python run_clef_gguf.py \
--base /path/to/Cloudflare/clef-flash \
--hsdump ./hsdump
The default record is the invoice example from the Cloudflare card (one
choice question and one noul question). --record file.json scores
another text record of the same shape. Hidden states are llama.cpp final-norm
embeddings (llama_set_embeddings_nextn, masked false), the read used by
Livesport/clef-flash-GGUF.
Limitations
- Text only. No images, no video.
- Post-training quantization. No recovery training.
- One calibration seed (1234). decision-v7 participated in the imatrix and the KL; treat that score as possibly high. transfer-v9 did not.
- No paired bootstrap and no confidence interval.
- llama.cpp chat or completion on this file is not a Clef decision. The head
is
joint_schema_model.py.
License and attribution
Derived from Cloudflare/clef-flash (© Cloudflare, Apache-2.0), itself post-trained from Qwen/Qwen3.5-9B (Apache-2.0). The hidden-state read follows Livesport/clef-flash-GGUF. Inference runs on llama.cpp.
繁中摘要
這是 Cloudflare/clef-flash 文字主幹的 measured-KL 混合精度 GGUF,3.300 GB(3,299,992,800 bytes)。decision-v7 正確率 0.8813(bf16 0.8861,保留 99.5%),transfer-v9 0.7859(bf16 0.8011,保留 98.1%)。單一 seed,沒有信賴區間。imatrix 與 KL 用了 decision-v7 的校正資料(imatrix 128 筆、seed 1234;KL 用其中前 32 筆),所以 decision-v7 可能偏樂觀;transfer-v9 沒有參與校正。沒有視覺。分類要跑 Python 的 JointSchemaHead,選項列取自原模型 bf16 lm_head;llama.cpp 單獨跑不出 clef 的判斷。output.weight 是 Q2_K,分類頭不讀它。norms、conv 在檔案裡是 F32,ssm_alpha / ssm_beta 是 BF16,都沒有再量化。
- Downloads last month
- 33
We're not able to determine the quantization variants.