clef
Derived from Cloudflare/clef. Weights revision 2f3de3dd. This represents the model implementation on Tenstorrent hardware. See the original model card for license, training, and evaluation details.
Clef is a 27B multimodal decision model from Cloudflare, post-trained from Qwen/Qwen3.8-27B. It reads one state (text, JSON, images or video) and a schema of typed questions (choice, score, noul) and returns a probability for every allowed option of every question in one forward pass, without generating text. This package serves it on two Tenstorrent Blackhole chips (tensor parallel 2) through the model's own FastAPI server, which implements the Jev/SystemOne API (POST /v1/systemone), so SystemOne clients work against it unchanged. The weights pointer above covers everything the server loads: the backbone shards, the vision encoder, the joint schema head (joint_head.safetensors), the tokenizer and processor configuration, and the author's joint_schema_model.py all live in the Cloudflare/clef repo at the pinned revision.
Runs on p150x2 (mesh P150x2).
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
At a glance
| Architecture | Qwen3_5ForConditionalGeneration backbone, 64 layers (48 Gated DeltaNet and 16 full-attention), hidden 5120, with a 27-block Qwen3.5 vision tower on device and a 65M-parameter joint schema head on the host CPU; prefill only, no decode |
| Hardware | p150x2 |
| License | apache-2.0 |
| Status | Experimental community bring-up |
Intended use
Direct use: Typed decisions over a state of up to 16,384 tokens, with or without images: classification, routing, triage, intent detection, extraction choices, policy and eligibility checks, multiple-choice reading comprehension, judging a proposed answer against stated criteria, and matching a caption or a label to an image. Workflows that act on confidence, with thresholds frozen on a labelled sample of the user's own workload.
Out-of-scope use: Text generation, chat, summarisation or open-ended question answering; the model only scores the options it is given. Fully automated decisions with legal, medical, financial, employment or similar consequences for people, without human review. The SystemOne permute and separate routes are not served. Languages other than English are not evaluated.
Quickstart
uv tool install tenstorrent # once — the Tenstorrent CLI, `tt`
tt model pull tt-hous/clef
tt serve tt-hous/clef
tt model pull (or tt-model pull --with-weights) downloads the Docker image and the Cloudflare/clef weights at 2f3de3dd85f379784083b0814d997ab627200f0c (into your HF cache; they are not in the image). tt serve (or tt-model serve) starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.
Without tt-cli — tt-model alone does the whole job:
tt-model pull tt-hous/clef --with-weights
tt-model serve tt-hous/clef
Request shape
The server implements the Jev/SystemOne API. A request has model, state (any string or
JSON value), questions (question id to question), and optional images and videos.
A question has type (noul, choice or score), optional instructions, and
criteria (a mapping of option id to description for choice, a list of ordered option
descriptions for score, optional true and false descriptions for noul). The
server accepts each image as base64-encoded image bytes (a bare string or a
data:image/...;base64, URL) or as an http(s) URL, and each video as a list of frames in
the same encodings. No API key is checked unless the server was started with CLEF_API_KEY
set. The port is 8008 (tt-model serve ... --port 8008).
Text example (the Jev/SystemOne example of the upstream model card), with curl
curl -s http://127.0.0.1:8008/v1/systemone -H 'content-type: application/json' -d '{
"model": "clef",
"state": "Our checkout started returning errors and orders are blocked.",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle the message?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}},
"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"}
}
}'
Image example (the receipt example of the upstream model card), with curl
IMG=$(base64 -w0 receipt.jpg)
curl -s http://127.0.0.1:8008/v1/systemone -H 'content-type: application/json' -d "{
\"model\": \"clef\",
\"state\": {\"task\": \"Review the attached receipt.\"},
\"images\": [\"$IMG\"],
\"questions\": {\"legible\": {\"type\": \"noul\", \"instructions\": \"Is the receipt total legible?\"}}
}"
The same two requests from Python
import base64
import requests
URL = "http://127.0.0.1:8008/v1/systemone"
text = requests.post(URL, timeout=300, json={
"model": "clef",
"state": "Our checkout started returning errors and orders are blocked.",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle the message?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}},
"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"},
},
}).json()
print(text["answers"]["department"]["choice"], text["answers"]["urgency"]["score"], text["answers"]["outage"]["noul"])
with open("receipt.jpg", "rb") as f:
image_b64 = base64.b64encode(f.read()).decode()
image = requests.post(URL, timeout=300, json={
"model": "clef",
"state": {"task": "Review the attached receipt."},
"images": [image_b64],
"questions": {"legible": {"type": "noul", "instructions": "Is the receipt total legible?"}},
}).json()
print(image["answers"]["legible"]["noul"], image["usage"]["input_tokens"], image["latency_ms"])
Reference behaviour
The response body is the one the author's joint_schema_model.systemone() returns
(model, answers, usage) plus latency_ms. The author's code on CPU (bf16) answers the
text example above with department = technical (probability 0.9155), urgency score 1.82
(Today at 0.8651) and outage noul 0.8995, with 300 input tokens. This package answers it
with department = technical (0.9138), urgency score 1.83 (Today at 0.8706) and outage
noul 0.9011, the same 300 input tokens, in 270 ms of model time (latency_ms 269.7, eager engine, measured from the pushed image); the largest
probability difference over the three questions is 0.0055. The receipt image of the upstream
card is not public; the image path was checked with a New Yorker cartoon and five candidate
captions (330 input tokens): caption D at 0.6402 against 0.6475 on the CPU, 498 ms on a new
state and 237 ms when the state is in the prefix cache. The full parity numbers are in
Performance.
Using it
This package exposes the model's own HTTP server on port 20000, not an OpenAI-compatible chat API — its request and response shapes are the model's own. See the author's notes above for the payload it expects.
Routes
| Route | Body |
|---|---|
POST /v1/systemone |
{"model", "answers", "usage": {"input_tokens", "output_tokens"}, "latency_ms"}. Answer shapes follow the upstream systemone(): choice (choice, confidence, probabilities), noul (noul, the probability of true), score (score, confidence, legend, probabilities). |
GET /v1/models |
One card per served name (clef) with the weights revision, device, mesh plan, engine mode (eager), prefix_planner, the precision knobs, the limits and the prefix-cache counters. |
GET /health, GET /v1/health |
{"status": "ok", "workers", "queued"}. Never behind the API key. |
A state over CLEF_MAX_STATE tokens, or a schema over CLEF_MAX_TAIL tokens, returns 422
with the token count. A request whose questions have no criteria (other than noul) returns
422, as upstream systemone() raises. In the traced opt-in mode (below) an image grid outside
the warm list, or a request with more than one image or video, returns 422.
Server environment
Set these through the serve profile env of the manifest (tt-model serve tt-hous/clef --profile p150x2;
tt-model serve has no --env flag) or a bare docker run --env. The defaults below are the
shipped configuration (eager engine, prefix planner on), the one every number in Performance
was measured with.
| Variable | Default | Meaning |
|---|---|---|
HF_MODEL |
Cloudflare/clef (set by tt-model) |
Hub id or snapshot directory of the weights. CLEF_MODEL, when set, is read first; the server writes the resolved snapshot directory back into CLEF_MODEL. |
CLEF_REVISION |
2f3de3dd85f379784083b0814d997ab627200f0c |
Revision used when the model id has no @rev. |
HF_HUB_OFFLINE |
unset | 1 resolves Hub ids from the local cache only. |
CLEF_MESH_SHAPE |
1x2 (set by tt-model from the profile) |
Mesh the server opens. 1x2 is one TP=2 worker. With exactly two visible chips the server opens the 1x2 mesh directly under FABRIC_1D; with more it opens a 1x4 parent and takes the 1x2 submesh at CLEF_SUBMESH_OFFSET. 1x4 and 2x2 (two TP=2 workers) are the planned p150x4 profile, not in this build. |
MESH_DEVICE |
P150x2 (set by tt-model) |
The mesh SKU; used only when CLEF_MESH_SHAPE is unset. |
CLEF_PARENT_MESH |
unset (auto) | 1x4 (FABRIC_1D) or 2x2 (FABRIC_2D): forces the parent-plus-submesh open for a 1x2 target. Needs four visible chips. |
CLEF_SUBMESH_OFFSET |
0,0 |
Offset of the 1x2 submesh inside the parent. |
CLEF_MAX_STATE |
16384 |
Engine max_state_len and the 422 limit for the state piece (system prompt and media placeholder tokens included). The upstream encode_record limit. |
CLEF_MAX_TAIL |
4096 |
Engine max_tail_len and the 422 limit for the schema plus the suffix. |
CLEF_TRUNCATE_STATES |
0 |
1 keeps the first CLEF_MAX_STATE tokens instead of refusing and marks the response. |
CLEF_PREFIX_CACHE |
4 |
States kept in the prefix cache (KV pages plus the Gated DeltaNet state snapshot per slot), clamped by the engine's snapshot slots; 0 disables the cache. |
CLEF_PLANNER |
1 |
The prefix planner picks the cached prefix length so a cache miss takes the fewest prefill passes. 0 restores the fixed 128-aligned rule. Reported as prefix_planner in /v1/models. |
CLEF_TRACED |
0 |
0 is the shipped eager engine: it accepts any image grid and any number of images or videos per request. 1 is the traced opt-in for a deployment with a fixed set of image grids: it needs CLEF_VISION_WARM_GRID, refuses an image or video whose grid is not in that list and a request with more than one image or video (422), and reserves CLEF_TRACE_REGION of DRAM. It gains 0 to 3 percent of latency at this depth. |
CLEF_VISION_WARM_GRID |
unset (required when CLEF_TRACED=1) |
;-separated t,h,w patch grids the traced server warms before capture and then accepts, for example 1,16,20;1,22,38;1,26,36;1,28,36;1,28,38;1,40,50;2,16,20. Ignored in eager mode. |
CLEF_TRACE_REGION |
1073741824 when CLEF_TRACED=1, else 0 |
trace_region_size for the mesh open. |
QWEN_GDN_CONV |
fir when CLEF_TRACED=1; not read in eager mode |
Gated DeltaNet causal conv kernel of the traced body. kda changes the numerics (one argmax flip on the 16 reference records). |
CLEF_API_KEY |
unset | Bearer key for /v1/* (health routes stay open). |
CLEF_ALLOW_REMOTE_IMAGES |
1 |
0 refuses http(s) image URLs with 422. Fetches use a 10 s timeout and a 20 MiB cap. |
CLEF_MODEL_NAMES |
clef |
Comma-separated names listed by /v1/models and accepted in model. |
CLEF_WARMUP |
1 |
0 skips the warmup request per worker at startup. |
CLEF_PRECISION |
unset (selected) |
Precision profile of tt/precision_defaults.py: selected (the shipped knobs, equal to the stage 1 configuration) or stage1. The QWEN36_* knobs it sets can be overridden one by one. |
CLEF_FAKE_ENGINE |
0 |
1 opens no device and serves deterministic pseudo-random hidden rows through the real head, tokenizer and processor (API tests on a CPU). |
Browser demo and CLI
The package ships a static demo page and a command-line client in
code/models/autoports/cloudflare_clef/demo/ of this repo (inside the image at
/opt/tt-metal/models/autoports/cloudflare_clef/demo). The page sends POST /v1/systemone
from the browser and draws every option probability: a support-triage preset from the
announcement, the invoice JSON state from the upstream card, a repeated-state preset that shows
the prefix cache, two New Yorker cartoons and a synthetic receipt as image requests, and a
replay of 24 labelled records (ARC-Challenge, BANKING77, New Yorker) with running accuracy.
Against this image on two chips every preset answers (8 of 8) and the replay scores 22 of 24.
With the server up on port 8008:
hf download tt-hous/clef --include "code/models/autoports/cloudflare_clef/demo/*" --local-dir clef-pkg
cd clef-pkg/code/models/autoports/cloudflare_clef/demo
DEMO_BIND=0.0.0.0 ./run.sh http://<server-host>:8008
Open http://<demo-host>:8080/#server=http://<server-host>:8008. The page only needs plain
HTML, CSS and JavaScript; fonts come from Google Fonts with system fallbacks. The same
directory is inside the image, so it can be served from there without a download:
docker run --rm -p 8080:8080 --entrypoint python3 $(docker images -q tt-model/clef | head -1) \
-m http.server 8080 --bind 0.0.0.0 --directory /opt/tt-metal/models/autoports/cloudflare_clef/demo
The CLI needs Python 3 and nothing else: python3 quickstart.py --all sends every preset,
python3 quickstart.py --replay runs the labelled replay, --preset "<title>" sends one, and
--show-request prints the request bodies without sending them. --server points it at
another host. Presets live in presets.json, shared by the page and the CLI; the
demo README in the same directory documents every panel and the measured numbers.
Expected performance
Model time per request (new / repeated state) and requests per second at 64 concurrent
clients, in the format of the kev model card, measured on 2026 Oct 05 with the host-side
scripts/serving_bench_remote.py against the shipped server configuration (eager engine,
prefix planner on) on a p300c box: two Blackhole chips of one p300 board as the 1x2 mesh.
| Device | 6 questions, short state | 5 questions, 2,200-token state | Requests/s, 64 clients |
|---|---|---|---|
| P150x2 (2 chips, TP=2), eager engine with the prefix planner [1] | 412.6 / 410.4 ms | 1,231.2 / 533.8 ms | 2.36 |
Startup: 106 s from the server's first log line to Application startup complete (mesh open
4 s, weights converted from the safetensors and the model built in 100.6 s, one warmup
request); no tensor cache is written, so every start pays it.
- New / repeated state: the first number of each pair is a request whose state the server has not seen (prefix-cache miss: the state prefix is prefilled, then the schema tail runs); the second is a request whose state is in the prefix cache (hit: the KV pages and the Gated DeltaNet state of the slot are restored, then the schema tail runs). The worker holds
CLEF_PREFIX_CACHE= 4 states. A state under 128 tokens has no cached prefix, so new and repeated are equal by construction in the short-state column (that state is 122 tokens with the system prompt). - Latency is the server's
latency_msfield: the wall time of the model section of the request on its worker thread (state prefill or cache restore, schema continuation, vision tower when images are present, joint schema head on the host in fp32), without queue wait and HTTP. Clef decides every question of a request jointly in one pass, so one request runs on one TP group and there is no per-question fan-out. - Precision: the
selectedprofile oftt/precision_defaults.py, which the datatype sweep set equal to the stage 1 configuration: MLP gate / up weights bfp8, MLP down bf16, attention and Gated DeltaNet projections bfp8, MLP matmuls HiFi2 with fp32 accumulation, Gated DeltaNet decay gate fp32, fp32 Gated DeltaNet recurrent state, bf16 KV cache, bf16 embeddings, vision tower bf16 weights and activations with fp32 accumulation, joint schema head fp32 on the host. - Code: tt-metal commit
c72c144e6e4(branchhous/clef-bringup, the tree the image was built from) on top ofd76d41fb52c(the Kev bring-up head) over7eac776e926(origin/main 2026 Oct 01). Every code path shipped in this package is byte-identical to the stage 3 serving commite75b184fc69, which produced every number above; the commits between them add only documentation, evaluation scripts and the demo directory (shipped undercode/but not imported by the server). - Method: port of kev's
scripts/serving_bench.pyover HTTP. Latency columns are the median of 20 requests after 2 warm-up requests. Throughput is 256 requests of the 6-question short-state case over 64 distinct states per client level (1, 8, 32, 64 clients), two passes, the second timed; it is flat in the client count (2.35 to 2.36 req/s) because one TP group serves one request at a time, and the queue wait grows linearly with the clients. 64 distinct 2,200-token states give 0.81 req/s. - [1] One worker on the 1x2 mesh, tensor parallel 2. Latency detail (new / repeated, p50 ms): 3 questions, 346-token blog triage state 268.2 / 265.0; 2 questions, short state 267.9 / 257.8; 5 questions, 370-token state 426.0 / 423.1. A request costs about 220 ms of fixed work plus about 0.35 s per 1,024-token chunk of prefill; a 16,384-token state (the limit) takes 7,471.8 ms on a miss and 1,600.0 ms on a hit.
Image requests
One image record (a New Yorker cartoon, 320 patches, with five candidate captions, 330 input
tokens): 497.7 ms on a new state and 237.1 ms on a repeated state; median 551 ms over the 8
reference image records (320 to 2,000 patches). Vision tower parity against the HF
model.visual on CPU on the 4 reference images: merged-output PCC 0.99880, 0.99822, 0.99909,
0.99898 (bar 0.97); blocks 0, 12, 23 at 0.99856 or better (bar 0.99) and block 26 at 0.99694
or better (bar 0.85); the 4-frame video record 0.99839. The largest image run end to end is
57,344 patches ((1, 224, 256), 3584 x 4096 pixels, 14,336 image tokens): 33.2 s per request
warm; a 2048 x 2048 image (16,384 patches, 4,096 tokens) takes 5.7 s. Above 4,096 patches
the tower time grows about quadratically. The first image of a new patch grid compiles its
programs once per box: 1 to 2 s at 2,048 padded rows, 6 to 9 s at 4,096, up to 26 s at
57,344.
Evaluations
Public benchmarks from the upstream model card, rendered once to SystemOne requests (no
prompt tuning, no option dropping) and run on this server; a CPU bf16 control (the author's
own joint_schema_model.py, same rendering) on a stratified 100-item sample of each
benchmark separates the rendering from the device precision. The Decision Index request data
behind the card numbers is not public, so the card numbers are indicative; this package does
not reproduce the Decision Index and does not claim to.
| Benchmark (test split) | n | Metric | This package, full test set | 100-item sample: this package / CPU bf16 control | Model card |
|---|---|---|---|---|---|
| ARC-Challenge | 1,172 | accuracy | 97.8 | 96.0 / 95.0 | 97.7 |
| BANKING77 | 3,080 | macro-F1 over 77 intents | 94.4 | 91.1 / 91.1 | 94.2 |
| New Yorker caption matching (image, 5 captions) | 528 | accuracy | 58.3 | 60.0 / 60.0 | 69.5 |
Coverage is complete on every file (0 rejected, 0 truncated, 0 unanswered). The 100-item samples carry a binomial 95 percent interval of about plus or minus 4 points on ARC and plus or minus 10 on New Yorker, so the sample column compares the device to the CPU control, not to the card. On New Yorker the device and the CPU control agree on the sample (60.0 / 60.0), so the 11-point gap to the card is the rendering of the public dataset to a SystemOne request against the Decision Index rendering, which is not public, and not the device. On ARC the device is one item above the control (+1.0 pp, one argmax flip at reference margin 0.37 out of 100); on BANKING77 the two agree (0 flips at margin, 2 near ties).
Kev's labelled suites (hard-v1, devtools-v1, documents-v1 test splits, SystemOne
requests, kev.benchmark --remote unchanged, clean block): accuracy, Brier and ECE
(expected calibration error, 10 bins on the top probability). Coverage complete (every
record evaluated, 0 rejected, 0 truncated). These suites are Kev-9B's own development
distribution; the Kev column is jaredpalmer/kev-9b served by its own port on one P150 of
the same box (tt-hous/kev-9b), two decision models on the same silicon and the same
labelled requests, not a ranking of the ports.
| Suite (test split) | questions | accuracy | Brier | ECE | Kev-9B on one P150: accuracy / ECE |
|---|---|---|---|---|---|
| hard-v1 | 1,088 | 0.774 | 0.321 | 0.044 | 0.826 / 0.053 |
| devtools-v1 | 1,073 | 0.741 | 0.399 | 0.143 | 0.787 / 0.101 |
| documents-v1 | 936 | 0.893 | 0.169 | 0.044 | 0.896 / 0.015 |
Parity against the author's CPU reference
Over HTTP against the served package (max |dp| is the largest absolute difference of an option probability; flips are argmax changes at a reference top-2 margin of at least 0.05):
| Set | questions | max dp | flips at margin | reference |
|---|---|---|---|---|
| 16 reference text records | 27 | 0.072 (mean 0.011) | 0 | CPU bf16 |
| same 16 records | 27 | 0.076 (mean 0.010) | 0 | CPU fp32 |
| 8 reference image records | 8 | 0.046 (mean 0.020) | 0 | CPU bf16 (0.038 against fp32) |
| 100 BANKING77 rows (1,824 to 1,871 tokens) | 100 | 0.110 | 0 (2 near ties) | CPU bf16 |
| 8 long records, 2,637 to 16,384 tokens | 32 | 0.056 | 0 | CPU bf16 |
| 64 development text records (dev64) | 95 | 0.192 | 1 (margin 0.17, dp 0.10) | CPU bf16 |
| 16 development image records (dev16) | 16 | 0.175 | 0 | CPU bf16 |
The eager engine's probabilities equal the traced engine's bit for bit; the served probabilities equal the engine's to the 4-decimal rounding of the API. The 16 reference records, the 8 image records and the 8 long records meet the bring-up bars (max dp at most 0.10, 0 flips at margin); dev64 and dev16 do not on the max-dp bar, see Limitations.
Limitations
- Hidden-state agreement with the HF reference is gated per layer, not per row: the teacher-forced per-layer PCC of every Gated DeltaNet layer is in the band of the attention layers (means 0.99972 to 0.99979 at 300 and 8,192 tokens, every layer except the last above 0.999). The whole-row PCC of the final hidden state against HF bf16 at 1,500 to 8,192 tokens is reported, not gated (all-rows mean 0.977 to 0.990; worst sampled row 0.76 at 1,500 tokens); HF bf16 itself fails a whole-row 0.99 bar against HF fp32 on 3.9 to 7.4 percent of rows, and the device is 3.5 to 5.9 times more often below that bar. The decision-level parity above is the gate that passes.
- Two precision-sensitive near-decision records: on a 200-record development set the device agrees with the CPU bf16 reference on 274 of 278 argmaxes and is within 0.36 pp in accuracy, but two questions (
hard-v1/multi_hop/development/00017 paged, dp 0.33, device wrong;hard-v1/probability/development/00084 value, dp 0.47, both wrong) move by up to 0.3 in probability while the reference is confident. No precision knob of the datatype sweep brings both into the normal band; a per-layer probe classifies them as precision sensitivity of a near-decision record (every layer inside the band above), not an op bug. The dev64 flip above (margin 0.17) is of the same kind. Expect a few such records per thousand on confident decisions. - Image evaluation: New Yorker caption matching is 58.3 against 69.5 on the card with the device equal to the CPU control on the sample, so the gap is the rendering of the public dataset, not the device. One development image record (
7b7ef383, dev16) is at max dp 0.175 with the argmax kept (reference margin 0.36); the other 15 are at 0.052 or below. The vision tower runs in bf16 with fp32 accumulation inside theqwen36port; its upstream bfp8 configuration puts the 8 reference records at max dp 0.22. - One TP group: requests are served one at a time, throughput is 2.36 req/s for short states and 0.81 req/s for new 2,200-token states at any client count, and the queue wait grows linearly with the clients (27 s at 64 clients). A second TP=2 worker on four chips (profile
p150x4) is planned and not in this build. - Startup is 106 s on every start: the weights convert from the safetensors each time (no tensor cache).
- States are limited to
CLEF_MAX_STATEtokens (16,384 by default, the upstreamencode_recordlimit) and schemas toCLEF_MAX_TAIL(4,096); longer requests return 422 instead of the upstream silent truncation (CLEF_TRUNCATE_STATES=1opts into truncation). The context the HF config advertises (262,144) is not served. - Images run through the on-device vision tower; the largest image validated end to end is 57,344 patches (3584 x 4096 pixels, 14,336 image tokens, 33 s per request) and the limit above it is the 16,384-token request budget, not the device. A request with many or large images is slow (about quadratic in patches above 4,096). Videos and images in one request are not supported together (one media kind per request).
- Per-grid first-request compile in eager mode: the first image of a patch grid the box has not seen compiles its programs once (1 to 2 s at up to 2,048 padded rows, 6 to 9 s at 4,096, 26 s at 57,344); the kernel cache under
~/.cache/tt-model/clef/cachepersists across container restarts, so this is a per-box cost. - Tracing is opt-in and constrained:
CLEF_TRACED=1needsCLEF_VISION_WARM_GRID, refuses image grids outside that list and more than one image or video per request (422), reserves a 1 GiB trace region and gains 0 to 3 percent. - The prefix cache holds
CLEF_PREFIX_CACHEstates (4 at the 16,384-token limit); only the 128-aligned state prefix is cached, the schema always runs, and a state under 128 tokens is never cached. The joint schema head runs on the host CPU in fp32 over every row of the record, so the host part grows with the state length (a 16,384-token cache hit is 1,600 ms). - The direct 1x2 mesh open with exactly two visible chips is the path this profile takes inside its container; the bring-up box ran every measurement through the four-chip parent-plus-submesh path (the same chips, the same TP=2 worker), so the direct open is exercised for the first time by the package validation.
- English only (the card's benchmarks are English; other languages are not evaluated here). Knowledge is set by the base model. Option order can change an answer.
Risks and safety considerations
Calibrated probabilities can create unwarranted trust; measure accuracy and calibration on a labelled sample of your own data before setting thresholds, and monitor production error rates. The model card's benchmark numbers come from non-public Decision Index renderings and are indicative only; this package does not reproduce them. A small fraction of confident decisions moves by up to 0.3 in probability between the device and the CPU reference (see Limitations). The server is open unless CLEF_API_KEY is set and accepts http(s) image URLs unless CLEF_ALLOW_REMOTE_IMAGES=0; states, images and videos may contain personal or confidential data, so apply your own access control and do not expose the server to untrusted networks without a proxy.
Licensing
Cloudflare/clef (backbone, vision encoder, joint schema head and joint_schema_model.py) is released under Apache-2.0; its model card states the base model Qwen/Qwen3.8-27B is Apache-2.0. The serving code under code/ is tt-metal (Apache-2.0) plus the Clef port under models/autoports/cloudflare_clef (Apache-2.0, see its NOTICE), which vendors the joint schema head classes from joint_schema_model.py and reuses the serving design of the Apache-2.0 kev project. The port's browser demo (in code/models/autoports/cloudflare_clef/demo, also inside the image) ships ten New Yorker caption contest cartoons from the jmhessel/newyorker_caption_contest dataset under CC BY 4.0, as its NOTICE states.
Related packages
Upstream model card: https://huggingface.co/Cloudflare/clef. Announcement: https://blog.cloudflare.com/clef-decision-models. Decision Index leaderboard: https://clef-evals.workers-ai-mle.workers.dev. The smaller variant Cloudflare/clef-flash is not packaged here. The same serving pattern on one Blackhole chip: tt-hous/kev-9b.
Feedback
Questions or problems with this package: open a discussion at https://huggingface.co/tt-hous/clef/discussions — that is what reaches its author. A problem with the tt tooling itself: tt report issue (collects your environment and opens a prefilled issue against tenstorrent/tt-cli). Product feedback: support@tenstorrent.com.
Provenance
The exact sources the image was built from — code/ in this repo is byte-identical to the model code inside the image:
| component | built from |
|---|---|
| tt-metal | c72c144e6e4d4cb1a29ee827cfb98d66b9883c30 (dirty tree — the image includes uncommitted changes) |
code/ digest |
f2f49abe472ac6d3 (sha256, first 16 hex digits) |
| built | 2026-10-06T14:13:56+00:00 by tt-model 0.1.0 |