kenning-large-v0.4
Kenning is a System One decision model: given program state and typed questions (noul yes/no, choice, score) it returns typed answers with calibrated probabilities in one forward pass. It does not generate text. It was built with SystemOne Builder.
Model
- Base model:
MoritzLaurer/deberta-v3-large-zeroshot-v2.0-c(MIT; fine-tuned on MNLI, FEVER-NLI (CC-BY-SA-3.0) and Mixtral-generated synthetic data (no non-commercial data)) - Architecture: cross-encoder; every candidate answer is scored against the state, and each question's distribution is
softmax(scores / T)with per-type temperatures fitted on held-out data - Temperatures:
{"score": 1.1618, "choice": 1.2214, "noul": 1.2214} - Max tokens per (state, answer) pair: 512
- Created: 2026-10-03T06:41:23Z
Training data
| Source | Rows | Licence |
|---|---|---|
fancyzhx/amazon_polarity |
2000 | Apache-2.0 |
clinc/clinc_oos |
3000 | CC-BY-3.0 |
nyu-mll/multi_nli |
20000 | OANC (permissive); fiction genre excluded |
google/civil_comments |
2000 | CC0-1.0 |
synthetic: written by Qwen/Qwen2.5-7B-Instruct from labelled scenarios |
1985 | generated (teacher: Qwen2.5-7B-Instruct, Apache-2.0) |
synthetic: 10 label-conditioned tasks written by Qwen/Qwen2.5-7B-Instruct |
2793 | generated (teacher: Qwen2.5-7B-Instruct, Apache-2.0) |
Distilled: every label was blended with the soft labels (probabilities) of Cloudflare/clef-flash (Apache-2.0); the target is 50% original label + 50% teacher.
Held-out results (split of the training pool)
| accuracy | Brier | ECE | |
|---|---|---|---|
| zero-shot (before training) | 0.670 | 0.409 | 0.182 |
| trained | 0.927 | 0.040 | 0.025 |
| trained + calibrated | 0.927 | 0.038 | 0.049 |
These are in-distribution numbers. Benchmark on your own data (and recalibrate on a few hundred labelled examples from it with systemone calibrate) before letting the model act automatically.
General benchmark (headline)
systemone bench --suite general: 1,328 held-out items, 30 questions in 7 families of state. Labels come from public held-out datasets, or from exact rules for generated records, logs and agent steps; no model labels them. The score is macro accuracy (the mean over questions). Both models ran on one RTX 3090, 5/5 identical on repeat.
| Family | kenning-large-v0.4 | Cloudflare clef-flash |
|---|---|---|
| All 30 questions | 0.625 | 0.813 |
| Text: evidence, sentiment, toxicity, prompt injection, intent, topic (14) | 0.754 | 0.841 |
| Conversations (1) | 0.958 | 1.000 |
| Answer quality: helpful, correct (2) | 0.377 | 0.447 |
| Agent decisions: tool call matches, should call a function, task completed (3) | 0.553 | 0.793 |
| Records: rules over JSON with numbers and dates (7) | 0.535 | 0.857 |
| Tables: is a statement true (1) | 0.480 | 0.860 |
| Logs: is a service failing, which one (2) | 0.290 | 0.740 |
| Latency p50 | 34 ms | 149 ms |
Kenning is close to Clef on text, about 20 times smaller and 4 times faster. On structured state (records, tables, agent steps, logs) it is far behind and often near chance. If your state looks like that, measure it on your own data first, or train it for your problem. Per-task results and method: Kenning docs.
Email suites
Recorded with systemone bench on suites never used for training. Automated is the share of items decided without a person (p >= 0.9 or <= 0.1); a threat auto-closed is a positive item the model was sure was negative.
| Suite | Items | Accuracy | Brier | ECE | Automated | Threats auto-closed | Latency p50 |
|---|---|---|---|---|---|---|---|
| Modern emails 2 (held out) | 20 | 0.750 | 0.151 | 0.186 | 20% | 0 | 47 ms |
| Modern emails | 20 | 1.000 | 0.024 | 0.106 | 70% | 0 | 49 ms |
| Phishing dataset | 50 | 0.780 / 0.820 | 0.151 | 0.211 | 44% | 0 | 88 ms |
| Layouts (4 state shapes) | 200 | 0.780 / 0.760 / 0.760 / 0.760 | 0.151 | 0.211 | 38% | 0 | 58 ms |
| Out of domain | 180 | 0.917 / 0.583 / 0.867 | 0.086 | 0.172 | 30% | 1 | 48 ms |
Where a suite asks several questions, accuracy lists each (out of domain: spam / emotion / news topic; layouts: one per layout). Comparisons with Cloudflare Clef and TypeSafe Jev on the same suites are in the Kenning docs.
Use
pip install "systemone-client[local]"
from systemone import Kenning, Noul, Choice
model = Kenning.from_pretrained("systemonedev/kenning-large-v0.4") # in-process, runs on CPU or GPU
answer = model.system_one(state={"ticket": "I was charged twice."},
questions={"billing": Noul("Is this about billing?")})
print(answer.nouls["billing"].noul)
Or serve it with SystemOne Builder's kenning service and call POST /v1/systemone with systemone.Client.
Limitations
- Structured state is the weak spot. Rules over records, numbers and dates, tables, agent steps and logs score 0.29–0.55 on the general benchmark (Clef: 0.74–0.86). Its training data was mostly short text.
- 512 tokens per (state, answer) pair. Longer states are truncated, so a decision that depends on the end of a long log or document is not seen.
- Calibration was fitted on the training distribution; probabilities on other data are scores until recalibrated (
systemone calibraterefits the temperatures on your labelled data). - Answer-quality judgements (is an answer helpful or correct) are hard for it and for Clef.
- Determinism: identical requests give identical answers on the same hardware and software; results can differ across GPUs or library versions.
Licence
The weights are released under the Apache License 2.0 (LICENSE). Upstream licences of the base model and of every training source are listed in NOTICE.md; some are share-alike (CC-BY-SA-3.0), so keep NOTICE.md with the weights.
Kenning implements a wire format compatible with TypeSafe AI's System One API. It is not affiliated with or endorsed by TypeSafe AI, and was not trained on TypeSafe outputs.
- Downloads last month
- 63