kev-9b

Kev-9B is a decision model: a LoRA adapter and a pointer head on Qwen3.5-9B-Base. It reads one document (the state) and a set of typed questions (choice, score, noul) and returns a calibrated probability distribution over the options of each question in one forward pass, without generating text. This package serves it on Tenstorrent Blackhole through the model's own FastAPI server, which implements TypeSafe's System One API (POST /v1/systemone), so the TypeSafe SDK works against it unchanged.

Runs on p150 or p300x2 โ€” see the serve profiles below.

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

At a glance

Architecture Qwen3.5-9B-Base backbone (24 Gated DeltaNet and 8 full-attention layers), merged LoRA adapter, pointer head; prefill only
Hardware p150, p300x2
License apache-2.0
Status Experimental community bring-up

Intended use

Direct use: Typed decisions over English documents of a few thousand tokens: classification, routing, triage, extraction choices, policy and eligibility checks, and judging a proposed answer against stated criteria. Workflows that act on confidence, with thresholds frozen on a labelled sample of the user's own workload.

Out-of-scope use: Text generation, chat, summarisation or open-ended question answering; the model only scores the options it is given. Fully automated decisions with legal, medical, financial, employment or similar consequences for people, without human review. Languages other than English.

Quickstart

uv tool install tenstorrent   # once โ€” the Tenstorrent CLI, `tt`
tt model pull tt-hous/kev-9b
tt serve tt-hous/kev-9b

tt model pull (or tt-model pull --with-weights) downloads the Docker image and the Qwen/Qwen3.5-9B-Base weights at 68c46c4b3498877f3ef123c856ecfde50c39f404 (into your HF cache; they are not in the image). tt serve (or tt-model serve) starts the model's own HTTP server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.

Without tt-cli โ€” tt-model alone does the whole job:

tt-model pull  tt-hous/kev-9b --with-weights
tt-model serve tt-hous/kev-9b

Client

The server implements TypeSafe's System One API. Install the SDK (pip install typesafe-sdk) and point it at the port this package publishes (8008 unless tt serve ... --port moved it). No API key is checked unless the server was started with KEV_API_KEY set; any string works.

from typesafe_sdk import Choice, Noul, TypeSafeClient

client = TypeSafeClient(api_key="local", base_url="http://127.0.0.1:8008", model="kev-latest")
response = client.system_one(
    state="I was charged twice for order 1182. Please refund one of the charges.",
    questions={
        "team": Choice(instructions="Which team should handle this?",
                       criteria={"billing": "Charges and refunds", "shipping": "Deliveries", "returns": "Exchanges"}),
        "urgent": Noul(instructions="Does this need a reply today?"),
    },
)
print(response.choices["team"].choice, response.nouls["urgent"].noul)

The same request with curl:

curl -s http://127.0.0.1:8008/v1/systemone -H 'content-type: application/json' -d '{
  "model": "kev-latest",
  "state": "I was charged twice for order 1182. Please refund one of the charges.",
  "questions": {
    "team": {"type": "choice", "instructions": "Which team should handle this?",
             "criteria": {"billing": "Charges and refunds", "shipping": "Deliveries", "returns": "Exchanges"}},
    "urgent": {"type": "noul", "instructions": "Does this need a reply today?"}
  }
}'

Demo

An optional browser and command-line demo lives in the tt-metal branch that holds this port (models/autoports/jaredpalmer_kev_9b/demo/ on github.com/housTT/tt-metal, branch hous/kev-9b-bringup). It is a static page plus a Python script that need nothing but Python 3; they are not part of this image.

B=https://raw.githubusercontent.com/housTT/tt-metal/hous/kev-9b-bringup/models/autoports/jaredpalmer_kev_9b/demo
mkdir -p kev-demo && cd kev-demo
for f in index.html presets.json quickstart.py run.sh README.md; do curl -fsSLO "$B/$f"; done
chmod +x run.sh
./run.sh http://127.0.0.1:8008        # prints the URL to open, DEMO_PORT=8081 if 8080 is taken
python3 quickstart.py --preset "Support ticket triage"
python3 quickstart.py --replay        # 24 labelled records, prints accuracy and latency

The page sends the presets to this server from the browser (the server allows any origin), draws one probability bar per option, and shows the server's latency_ms next to the client wall time. Presets: a support-ticket triage with six questions, a 2,200-token document whose second request shows the prefix cache (1,533 ms, then 527 ms on one chip), complaint routing, answer checking, code review and news topic records with their labels, and a 24-record replay with a running accuracy. Measured on one P150 on 2026 Oct 02 against this package: support ticket 606 ms, replay 33 of 41 questions correct in 8.4 s. With KEV_API_KEY set the page cannot reach the server from another origin (the preflight is refused); the CLI works with --api-key.

Adapter weights

weights above is the base checkpoint. The LoRA adapter and pointer head come from a second Hub repo, jaredpalmer/kev-9b at revision db029f08b290afd9fee4aa4bbcd9ae48602d1eb0 (KEV_RUN in the serve env). The server downloads it into the mounted HF cache at first boot. To stage it before an offline boot: hf download jaredpalmer/kev-9b --revision db029f08b290afd9fee4aa4bbcd9ae48602d1eb0.

Serve profiles

One image serves every profile below; pick one with --profile.

profile hardware mesh
p150 (default) p150 P150
p300x2 p300x2 QB2

Using it

This package exposes the model's own HTTP server on port 20000, not an OpenAI-compatible chat API โ€” its request and response shapes are the model's own. See the author's notes above for the payload it expects.

Routes

Route Body
POST /v1/systemone {"model", "answers", "usage": {"input_tokens", "output_tokens"}, "latency_ms"}. Answer shapes follow kev: choice (choice, confidence, probabilities), noul (noul), score (score, legend, probabilities, confidence).
GET /v1/models One card per served name (kev-latest, jev-latest) with the run, base, device, backend, temperature, limits and per-worker prefix-cache counters.
GET /health, GET /v1/health {"status": "ok", "workers", "queued"}. Never behind the API key.

Every response carries x-typesafe-request-id (echoed from the request when given) and server-timing. A state over the limit returns 422 with the token count.

Server environment

Set these through the serve profile env (tt serve tt-hous/kev-9b --profile ...) or a bare docker run --env:

Variable Default Meaning
KEV_RUN jaredpalmer/kev-9b@db029f08b290afd9fee4aa4bbcd9ae48602d1eb0 Adapter and head.pt: Hub id with optional @revision, or a local directory.
KEV_MESH_SHAPE set by tt-model from the profile (1x1, 2x2) Parent mesh shape; one worker per 1x1 submesh.
KEV_DEVICES all Which submeshes get a worker, by index (0,2).
KEV_DEVICE_ID 0 Chip id for the single-chip (1x1) path.
KEV_PREFIX_CACHE 8 States kept per worker (clamped by the engine's snapshot slots).
KEV_FANOUT 1 1 lets idle workers each take a share of one request's questions; 0 sends every request to one worker. Under load the planner stops fanning out by itself.
KEV_FANOUT_BACKLOG_MS 200 A worker with more estimated backlog than this is not idle for fan-out.
KEV_PERF_SUMMARY shipped doc/optimized/perf_summary.json Cost model (per-block state and tail-bucket milliseconds) the dispatcher plans with.
KEV_PRECISION selected Weight precision profile: selected (bfp8 gate/up/down/proj, LoFi) or baseline (bfp4 gate/up).
KEV_MATMUL_POLICY 1 0 disables the per-shape matmul program-config policy.
KEV_TRACED 1 0 builds the eager engine without traces.
KEV_MAX_STATE 65536 Engine max_state_len and the 422 limit for states.
KEV_MAX_QUESTION 2048 422 limit for a question row.
KEV_TRUNCATE_STATES 0 1 reads the first KEV_MAX_STATE state tokens instead of refusing.
KEV_TRACE_REGION 1073741824 trace_region_size for the device open.
KEV_API_KEY unset Bearer key for /v1/* (health routes stay open).

Expected performance

Model time per request (new / repeated state) and requests per second at 64 concurrent clients, in the format of the upstream kev model card. Measured 2026 Oct 02 on a p300c box (four Blackhole P150 chips) with the host-side scripts/serving_bench_remote.py.

Device 6 questions, short state 5 questions, 2,200-token state Requests/s, 64 clients
P150 (1 chip) [1] 607.8 / 607.7 ms 1,533.5 / 527.0 ms 1.6
P150 x4 (data parallel, fan-out) [2] 203.9 / 203.9 ms 1,533.4 / 527.0 ms 6.5
  • New / repeated state: the first number of each pair is a request whose state the server has not seen (prefix-cache miss: the state is prefilled, then the questions run); the second is a request whose state is in the prefix cache (hit: the KV and GDN state are restored, then the questions run). Each worker holds 8 states.
  • Latency is the server's latency_ms field: the model time of the request in its worker thread (state prefill or cache restore, question rows, pointer head), without queue wait and HTTP. With fan-out it is the critical path over the workers that shared the request.
  • Precision: bfp8 weights for every projection including the MLP gate / up, LoFi matmuls with fp32 accumulation, bf16 activations and KV cache, bf16 GDN state. Prefill is traced; the matmul program-config policy is on.
  • Code: tt-metal commit ed917633df8 (branch hous/kev-9b-bringup) on top of 7eac776e926 (origin/main). This package ships that code.
  • Method: kev scripts/serving_bench.py semantics over HTTP. Latency columns are the median of 20 requests after 2 warm-up requests. Throughput is the full method: 256 development records of the 6-question short-state case (distinct states) plus 64 long-state requests per client level, two passes, the second timed, at 64 clients.
  • [1] One worker on one chip, whole-request dispatch (measured with KEV_FANOUT=0; with one worker the planner never fans out, so profile p150 with the default KEV_FANOUT=1 runs the same path).
  • [2] Four workers, one per chip, question-level fan-out (KEV_FANOUT=1, degrade threshold 200 ms of backlog). Profile p300x2 runs this configuration. The short-state request is split over the four chips; a 2,200-token state stays on one chip. Under load the planner falls back to whole requests, so the 64-client throughput equals the whole-request figure.

Evaluations

kev's kev.benchmark --remote on the four-chip server (test splits, run once, after the development splits). Metrics are kev's clean block: accuracy and ECE (expected calibration error, 10 bins on the top probability). Model-card numbers are the upstream fp32 evaluation.

Suite / test split n (questions) acc ECE model card acc / ECE
hard-v1 1,088 0.826 0.053 0.834 / 0.054
devtools-v1 1,073 0.788 0.101 0.791 / 0.098
documents-v1 936 0.896 0.015 0.900 / 0.017
hard-v1 + devtools-v1 audited (pooled) 1,861 0.815 0.060 0.822
breadth-v1 not run (private data) 0.698 / 0.034

Parity against kev's CPU fp32 predictor: on the 16 reference records (29 questions) max |dp| 0.088, mean 0.027, 1 argmax flip on a near tie (fp32 top-2 margin 0.03); on 64 development records (104 questions) max |dp| 0.416, mean 0.037, 2 flips at margin >= 0.05.

Known gaps: accuracy is 0.3 to 0.8 points below the card on every test split (the bfp8 / LoFi device precision; ECE matches within 0.003); breadth-v1 is not evaluated; long-state throughput is prefill-bound (64 distinct 2,200-token states give 2.5 req/s on four chips); the readback of the gathered hidden rows releases the GIL in this code because ttnn.from_device does not.

Limitations

  • The adapter is a second Hub pointer that tt-model does not manage: tt model pull fetches only the base checkpoint, and the first boot needs network (or a pre-populated HF cache) to fetch jaredpalmer/kev-9b.
  • States are limited to KEV_MAX_STATE tokens (65,536 by default, the upstream serving limit; the upstream model card validates 8,192). Longer states return 422 unless KEV_TRUNCATE_STATES=1.
  • The prefix cache holds KEV_PREFIX_CACHE states per worker (8 slots at the default state limit on one chip); only the 128-token-aligned prefix of a state is cached, so a hit on a short state saves little.
  • Requests are not batched across questions of different states; throughput scales with the number of workers (chips). Question fan-out across idle workers lowers single-request latency, not throughput.
  • English only. Knowledge is set by the base model. Option order can change an answer.
  • Accuracy is 0.3 to 0.8 points below the upstream fp32 model card on every test split (bfp8 / LoFi device precision); parity against fp32 on the 16 reference records is max |dp| 0.088 with one near-tie flip. See Performance.

Risks and safety considerations

Calibrated probabilities can create unwarranted trust; the temperature was fitted on public held-out datasets and does not transfer to every workload. Measure accuracy and calibration on a labelled sample of your own data before setting thresholds, and monitor production error rates. The server is open unless KEV_API_KEY is set. States may contain personal or confidential data; apply your own access control.

Licensing

Adapter and pointer head (jaredpalmer/kev-9b) are Apache-2.0; the upstream model card states the base model Qwen3.5-9B-Base is Apache-2.0. The serving code under code/ is tt-metal (Apache-2.0) plus kev code vendored under models/autoports/jaredpalmer_kev_9b (Apache-2.0, see its NOTICE).

Related packages

Upstream model card and serving reference: https://huggingface.co/jaredpalmer/kev-9b and https://github.com/jaredpalmer/kev.

Feedback

Questions or problems with this package: open a discussion at https://huggingface.co/tt-hous/kev-9b/discussions โ€” that is what reaches its author. A problem with the tt tooling itself: tt report issue (collects your environment and opens a prefilled issue against tenstorrent/tt-cli). Product feedback: support@tenstorrent.com.

Provenance

The exact sources the image was built from โ€” code/ in this repo is byte-identical to the model code inside the image:

component built from
tt-metal a local checkout โ€” commit not published
code/ digest 98f809a57e6c5d55 (sha256, first 16 hex digits)
built 2026-10-02T02:43:02+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for tt-hous/kev-9b

Finetuned
(611)
this model