kenning-large-v0.5
Kenning is a System One decision model: given program state and typed questions (noul yes/no, choice, score) it returns typed answers with calibrated probabilities in one forward pass. It does not generate text. Built with SystemOne Builder.
v0.5 redesigns the training data for structured state (records, tables, agent steps, logs), the families v0.4 was weakest on.
Model
- Base model:
MoritzLaurer/deberta-v3-large-zeroshot-v2.0-c(MIT; trained on MNLI, FEVER-NLI and Mixtral-generated synthetic data; no non-commercial data) - Architecture: 435M cross-encoder; every candidate answer is scored against the state, each question's distribution is
softmax(scores / T)with per-type temperatures fitted on held-out data - Temperatures:
{"noul": 2.34, "choice": 1.82, "score": 2.34} - Max tokens per (state, answer) pair: 512
- Deterministic: identical requests give identical answers on the same hardware/software
Training data
A balanced ~15k-row subset of the v0.5 corpus (permissively licensed or generated; every source listed in NOTICE.md):
- Public labelled datasets: MNLI, CLINC, Amazon reviews, Civil Comments, HelpSteer2 (answer quality), jailbreak-classification, prompt-injections, ROPES (passage reasoning). Train splits only; the benchmark uses the held-out splits.
- Rule-generated structured decisions (exact labels): expenses, shipping, leave, inventory, incident SLAs, eligibility, loans, subscriptions, tables, tool-call checks, agent traces, access/deploy logs.
- Teacher-written problem cases (blind-checked): computer-use, chat moderation, insurance claims, log triage, change review, support threads.
- Synthetic emails and tasks from an open LLM.
Not distilled from Clef this round (Clef's batch endpoint hangs on long inputs); text calibration is unchanged from v0.4.
Held-out results (split of the training pool)
| accuracy | Brier | ECE | |
|---|---|---|---|
| zero-shot (before training) | 0.587 | 0.560 | 0.146 |
| trained | 0.924 | 0.115 | 0.044 |
| trained + calibrated | 0.924 | 0.104 | 0.012 |
In-distribution numbers; they include easy exact-label rows. Benchmark on your own data (and recalibrate with systemone calibrate) before acting automatically.
General benchmark
systemone bench --suite general: 1,328 held-out items, 30 questions in 7 families. Accuracy per family vs Cloudflare Clef (clef-flash) and TypeSafe Jev (jev-latest), run for comparison only and never trained on. One RTX 3090 each.
| Family | v0.5 | Clef | Jev |
|---|---|---|---|
| Macro (30 questions) | 0.679 | 0.813 | 0.840 |
| Text (evidence, sentiment, toxicity, injection, intent, topic) | 0.756 | 0.841 | 0.834 |
| Conversation | 0.979 | 1.000 | 1.000 |
| Answer quality (helpful, correct) | 0.479 | 0.447 | 0.498 |
| Agent (tool calls, task completion) | 0.727 | 0.793 | 0.900 |
| Records (rules over JSON) | 0.587 | 0.857 | 0.921 |
| Tables | 0.520 | 0.860 | 0.940 |
| Logs | 0.520 | 0.740 | 0.720 |
| Latency p50 | 35 ms | 125 ms | 152 ms |
v0.5 is ~20x smaller than Clef and ~4x faster, beats both on answer quality, and made big gains over v0.4 (tool-call match 0.50→0.84, jailbreak 0.64→0.96, failing-service 0.08→0.52). It is still well behind on the decisions that need computation over the state — records, tables, logs — a limit of the cross-encoder architecture. That is the focus of the experimental Kenning-XL (v0.6), a decoder with deliberate reasoning that halves the macro gap to Clef (details).
Use
pip install "systemone-client[local]"
from systemone import Kenning, Noul
model = Kenning.from_pretrained("systemonedev/kenning-large-v0.5") # in-process, CPU or GPU
answer = model.system_one(state={"ticket": "I was charged twice."},
questions={"billing": Noul("Is this about billing?")})
print(answer.nouls["billing"].noul)
Or serve it with SystemOne Builder's kenning service and call POST /v1/systemone.
Limitations
- Structured state is the weak spot. Numeric records, tables and long logs score 0.52–0.59 (Clef 0.74–0.86): the cross-encoder scores each option in one short pass with nowhere to add numbers or scan a column.
- 512 tokens per (state, answer) pair. Longer state is truncated.
- Calibration is fitted on the training distribution; on other data the probabilities are scores until refitted (
systemone calibrate). - Determinism holds on the same hardware/software; results can differ across GPUs or library versions.
Licence
Weights under the Apache License 2.0 (LICENSE). Upstream licences of the base model and every training source are in NOTICE.md; some are share-alike, so keep NOTICE.md with the weights. Kenning implements a wire format compatible with TypeSafe AI's System One API; it is not affiliated with or endorsed by TypeSafe AI, and was not trained on TypeSafe outputs.
- Downloads last month
- 21