kev-9b
Kev-9B is a decision model: a LoRA adapter and a pointer head on Qwen3.5-9B-Base. It reads one document (the state) and a set of typed questions (choice, score, noul) and returns a calibrated probability distribution over the options of each question in one forward pass, without generating text. This package serves it on Tenstorrent Blackhole through the model's own FastAPI server, which implements TypeSafe's System One API (POST /v1/systemone), so the TypeSafe SDK works against it unchanged.
Runs on p150 or p300x2 โ see the serve profiles below.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
At a glance
| Architecture | Qwen3.5-9B-Base backbone (24 Gated DeltaNet and 8 full-attention layers), merged LoRA adapter, pointer head; prefill only |
| Hardware | p150, p300x2 |
| License | apache-2.0 |
| Status | Experimental community bring-up |
Intended use
Direct use: Typed decisions over English documents of a few thousand tokens: classification, routing, triage, extraction choices, policy and eligibility checks, and judging a proposed answer against stated criteria. Workflows that act on confidence, with thresholds frozen on a labelled sample of the user's own workload.
Out-of-scope use: Text generation, chat, summarisation or open-ended question answering; the model only scores the options it is given. Fully automated decisions with legal, medical, financial, employment or similar consequences for people, without human review. Languages other than English.
Quickstart
uv tool install tenstorrent # once โ the Tenstorrent CLI, `tt`
tt model pull tt-hous/kev-9b
tt serve tt-hous/kev-9b
tt model pull (or tt-model pull --with-weights) downloads the Docker image and the Qwen/Qwen3.5-9B-Base weights at 68c46c4b3498877f3ef123c856ecfde50c39f404 (into your HF cache; they are not in the image). tt serve (or tt-model serve) starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.
Without tt-cli โ tt-model alone does the whole job:
tt-model pull tt-hous/kev-9b --with-weights
tt-model serve tt-hous/kev-9b
Client
The server implements TypeSafe's System One API. Install the SDK (pip install typesafe-sdk)
and point it at the port this package publishes (8008 unless tt serve ... --port moved it).
No API key is checked unless the server was started with KEV_API_KEY set; any string works.
from typesafe_sdk import Choice, Noul, TypeSafeClient
client = TypeSafeClient(api_key="local", base_url="http://127.0.0.1:8008", model="kev-latest")
response = client.system_one(
state="I was charged twice for order 1182. Please refund one of the charges.",
questions={
"team": Choice(instructions="Which team should handle this?",
criteria={"billing": "Charges and refunds", "shipping": "Deliveries", "returns": "Exchanges"}),
"urgent": Noul(instructions="Does this need a reply today?"),
},
)
print(response.choices["team"].choice, response.nouls["urgent"].noul)
The same request with curl:
curl -s http://127.0.0.1:8008/v1/systemone -H 'content-type: application/json' -d '{
"model": "kev-latest",
"state": "I was charged twice for order 1182. Please refund one of the charges.",
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges and refunds", "shipping": "Deliveries", "returns": "Exchanges"}},
"urgent": {"type": "noul", "instructions": "Does this need a reply today?"}
}
}'
Demo
An optional browser and command-line demo lives in the tt-metal branch that holds this port
(models/autoports/jaredpalmer_kev_9b/demo/ on
github.com/housTT/tt-metal, branch hous/kev-9b-bringup).
It is a static page plus a Python script that need nothing but Python 3; they are not part of this image.
B=https://raw.githubusercontent.com/housTT/tt-metal/hous/kev-9b-bringup/models/autoports/jaredpalmer_kev_9b/demo
mkdir -p kev-demo && cd kev-demo
for f in index.html presets.json quickstart.py run.sh README.md; do curl -fsSLO "$B/$f"; done
chmod +x run.sh
./run.sh http://127.0.0.1:8008 # prints the URL to open, DEMO_PORT=8081 if 8080 is taken
python3 quickstart.py --preset "Support ticket triage"
python3 quickstart.py --replay # 24 labelled records, prints accuracy and latency
The page sends the presets to this server from the browser (the server allows any origin),
draws one probability bar per option, and shows the server's latency_ms next to the client
wall time. Presets: a support-ticket triage with six questions, a 2,200-token document whose
second request shows the prefix cache (1,533 ms, then 527 ms on one chip), complaint routing,
answer checking, code review and news topic records with their labels, and a 24-record replay
with a running accuracy. Measured on one P150 on 2026 Oct 02 against this package: support
ticket 606 ms, replay 33 of 41 questions correct in 8.4 s. With KEV_API_KEY set the page
cannot reach the server from another origin (the preflight is refused); the CLI works with
--api-key.
Adapter weights
weights above is the base checkpoint. The LoRA adapter and pointer head come from a second
Hub repo, jaredpalmer/kev-9b at revision db029f08b290afd9fee4aa4bbcd9ae48602d1eb0
(KEV_RUN in the serve env). The server downloads it into the mounted HF cache at first boot.
To stage it before an offline boot: hf download jaredpalmer/kev-9b --revision db029f08b290afd9fee4aa4bbcd9ae48602d1eb0.
Serve profiles
One image serves every profile below; pick one with --profile.
| profile | hardware | mesh |
|---|---|---|
p150 (default) |
p150 | P150 |
p300x2 |
p300x2 | QB2 |
Using it
This package exposes the model's own HTTP server on port 20000, not an OpenAI-compatible chat API โ its request and response shapes are the model's own. See the author's notes above for the payload it expects.
Routes
| Route | Body |
|---|---|
POST /v1/systemone |
{"model", "answers", "usage": {"input_tokens", "output_tokens"}, "latency_ms"}. Answer shapes follow kev: choice (choice, confidence, probabilities), noul (noul), score (score, legend, probabilities, confidence). |
GET /v1/models |
One card per served name (kev-latest, jev-latest) with the run, base, device, backend, temperature, limits and per-worker prefix-cache counters. |
GET /health, GET /v1/health |
{"status": "ok", "workers", "queued"}. Never behind the API key. |
Every response carries x-typesafe-request-id (echoed from the request when given) and server-timing.
A state over the limit returns 422 with the token count.
Server environment
Set these through the serve profile env (tt serve tt-hous/kev-9b --profile ...) or a bare docker run --env:
| Variable | Default | Meaning |
|---|---|---|
KEV_RUN |
jaredpalmer/kev-9b@db029f08b290afd9fee4aa4bbcd9ae48602d1eb0 |
Adapter and head.pt: Hub id with optional @revision, or a local directory. |
KEV_MESH_SHAPE |
set by tt-model from the profile (1x1, 2x2) |
Parent mesh shape; one worker per 1x1 submesh. |
KEV_DEVICES |
all |
Which submeshes get a worker, by index (0,2). |
KEV_DEVICE_ID |
0 |
Chip id for the single-chip (1x1) path. |
KEV_PREFIX_CACHE |
8 |
States kept per worker (clamped by the engine's snapshot slots). |
KEV_FANOUT |
1 |
1 lets idle workers each take a share of one request's questions; 0 sends every request to one worker. Under load the planner stops fanning out by itself. |
KEV_FANOUT_BACKLOG_MS |
200 |
A worker with more estimated backlog than this is not idle for fan-out. |
KEV_PERF_SUMMARY |
shipped doc/optimized/perf_summary.json |
Cost model (per-block state and tail-bucket milliseconds) the dispatcher plans with. |
KEV_PRECISION |
selected |
Weight precision profile: selected (bfp8 gate/up/down/proj, LoFi) or baseline (bfp4 gate/up). |
KEV_MATMUL_POLICY |
1 |
0 disables the per-shape matmul program-config policy. |
KEV_TRACED |
1 |
0 builds the eager engine without traces. |
KEV_MAX_STATE |
65536 |
Engine max_state_len and the 422 limit for states. |
KEV_MAX_QUESTION |
2048 |
422 limit for a question row. |
KEV_TRUNCATE_STATES |
0 |
1 reads the first KEV_MAX_STATE state tokens instead of refusing. |
KEV_TRACE_REGION |
1073741824 |
trace_region_size for the device open. |
KEV_API_KEY |
unset | Bearer key for /v1/* (health routes stay open). |
Expected performance
Model time per request (new / repeated state) and requests per second at 64 concurrent clients,
in the format of the upstream kev model card. Measured 2026 Oct 02 on a p300c box (four Blackhole
P150 chips) with the host-side scripts/serving_bench_remote.py.
| Device | 6 questions, short state | 5 questions, 2,200-token state | Requests/s, 64 clients |
|---|---|---|---|
| P150 (1 chip) [1] | 607.8 / 607.7 ms | 1,533.5 / 527.0 ms | 1.6 |
| P150 x4 (data parallel, fan-out) [2] | 203.9 / 203.9 ms | 1,533.4 / 527.0 ms | 6.5 |
- New / repeated state: the first number of each pair is a request whose state the server has not seen (prefix-cache miss: the state is prefilled, then the questions run); the second is a request whose state is in the prefix cache (hit: the KV and GDN state are restored, then the questions run). Each worker holds 8 states.
- Latency is the server's
latency_msfield: the model time of the request in its worker thread (state prefill or cache restore, question rows, pointer head), without queue wait and HTTP. With fan-out it is the critical path over the workers that shared the request. - Precision: bfp8 weights for every projection including the MLP gate / up, LoFi matmuls with fp32 accumulation, bf16 activations and KV cache, bf16 GDN state. Prefill is traced; the matmul program-config policy is on.
- Code: tt-metal commit
ed917633df8(branchhous/kev-9b-bringup) on top of7eac776e926(origin/main). This package ships that code. - Method: kev
scripts/serving_bench.pysemantics over HTTP. Latency columns are the median of 20 requests after 2 warm-up requests. Throughput is the full method: 256 development records of the 6-question short-state case (distinct states) plus 64 long-state requests per client level, two passes, the second timed, at 64 clients. - [1] One worker on one chip, whole-request dispatch (measured with
KEV_FANOUT=0; with one worker the planner never fans out, so profilep150with the defaultKEV_FANOUT=1runs the same path). - [2] Four workers, one per chip, question-level fan-out (
KEV_FANOUT=1, degrade threshold 200 ms of backlog). Profilep300x2runs this configuration. The short-state request is split over the four chips; a 2,200-token state stays on one chip. Under load the planner falls back to whole requests, so the 64-client throughput equals the whole-request figure.
Evaluations
kev's kev.benchmark --remote on the four-chip server (test splits, run once, after the development splits). Metrics are kev's clean block: accuracy and ECE (expected calibration error, 10 bins on the top probability). Model-card numbers are the upstream fp32 evaluation.
| Suite / test split | n (questions) | acc | ECE | model card acc / ECE |
|---|---|---|---|---|
| hard-v1 | 1,088 | 0.826 | 0.053 | 0.834 / 0.054 |
| devtools-v1 | 1,073 | 0.788 | 0.101 | 0.791 / 0.098 |
| documents-v1 | 936 | 0.896 | 0.015 | 0.900 / 0.017 |
| hard-v1 + devtools-v1 audited (pooled) | 1,861 | 0.815 | 0.060 | 0.822 |
| breadth-v1 | not run (private data) | 0.698 / 0.034 |
Parity against kev's CPU fp32 predictor: on the 16 reference records (29 questions) max |dp| 0.088, mean 0.027, 1 argmax flip on a near tie (fp32 top-2 margin 0.03); on 64 development records (104 questions) max |dp| 0.416, mean 0.037, 2 flips at margin >= 0.05.
Known gaps: accuracy is 0.3 to 0.8 points below the card on every test split (the bfp8 / LoFi device precision; ECE matches within 0.003); breadth-v1 is not evaluated; long-state throughput is prefill-bound (64 distinct 2,200-token states give 2.5 req/s on four chips); the readback of the gathered hidden rows releases the GIL in this code because ttnn.from_device does not.
Limitations
- The adapter is a second Hub pointer that tt-model does not manage:
tt model pullfetches only the base checkpoint, and the first boot needs network (or a pre-populated HF cache) to fetchjaredpalmer/kev-9b. - States are limited to
KEV_MAX_STATEtokens (65,536 by default, the upstream serving limit; the upstream model card validates 8,192). Longer states return 422 unlessKEV_TRUNCATE_STATES=1. - The prefix cache holds
KEV_PREFIX_CACHEstates per worker (8 slots at the default state limit on one chip); only the 128-token-aligned prefix of a state is cached, so a hit on a short state saves little. - Requests are not batched across questions of different states; throughput scales with the number of workers (chips). Question fan-out across idle workers lowers single-request latency, not throughput.
- English only. Knowledge is set by the base model. Option order can change an answer.
- Accuracy is 0.3 to 0.8 points below the upstream fp32 model card on every test split (bfp8 / LoFi device precision); parity against fp32 on the 16 reference records is max |dp| 0.088 with one near-tie flip. See Performance.
Risks and safety considerations
Calibrated probabilities can create unwarranted trust; the temperature was fitted on public held-out datasets and does not transfer to every workload. Measure accuracy and calibration on a labelled sample of your own data before setting thresholds, and monitor production error rates. The server is open unless KEV_API_KEY is set. States may contain personal or confidential data; apply your own access control.
Licensing
Adapter and pointer head (jaredpalmer/kev-9b) are Apache-2.0; the upstream model card states the base model Qwen3.5-9B-Base is Apache-2.0. The serving code under code/ is tt-metal (Apache-2.0) plus kev code vendored under models/autoports/jaredpalmer_kev_9b (Apache-2.0, see its NOTICE).
Related packages
Upstream model card and serving reference: https://huggingface.co/jaredpalmer/kev-9b and https://github.com/jaredpalmer/kev.
Feedback
Questions or problems with this package: open a discussion at https://huggingface.co/tt-hous/kev-9b/discussions โ that is what reaches its author. A problem with the tt tooling itself: tt report issue (collects your environment and opens a prefilled issue against tenstorrent/tt-cli). Product feedback: support@tenstorrent.com.
Provenance
The exact sources the image was built from โ code/ in this repo is byte-identical to the model code inside the image:
| component | built from |
|---|---|
| tt-metal | a local checkout โ commit not published |
code/ digest |
98f809a57e6c5d55 (sha256, first 16 hex digits) |
| built | 2026-10-02T02:43:02+00:00 by tt-model 0.1.0 |
Model tree for tt-hous/kev-9b
Base model
Qwen/Qwen3.5-9B-Base