kenning-large-v0.4

Kenning is a System One decision model: given program state and typed questions (noul yes/no, choice, score) it returns typed answers with calibrated probabilities in one forward pass. It does not generate text. It was built with SystemOne Builder.

Model

  • Base model: MoritzLaurer/deberta-v3-large-zeroshot-v2.0-c (MIT; fine-tuned on MNLI, FEVER-NLI (CC-BY-SA-3.0) and Mixtral-generated synthetic data (no non-commercial data))
  • Architecture: cross-encoder; every candidate answer is scored against the state, and each question's distribution is softmax(scores / T) with per-type temperatures fitted on held-out data
  • Temperatures: {"score": 1.1618, "choice": 1.2214, "noul": 1.2214}
  • Max tokens per (state, answer) pair: 512
  • Created: 2026-10-03T06:41:23Z

Training data

Source Rows Licence
fancyzhx/amazon_polarity 2000 Apache-2.0
clinc/clinc_oos 3000 CC-BY-3.0
nyu-mll/multi_nli 20000 OANC (permissive); fiction genre excluded
google/civil_comments 2000 CC0-1.0
synthetic: written by Qwen/Qwen2.5-7B-Instruct from labelled scenarios 1985 generated (teacher: Qwen2.5-7B-Instruct, Apache-2.0)
synthetic: 10 label-conditioned tasks written by Qwen/Qwen2.5-7B-Instruct 2793 generated (teacher: Qwen2.5-7B-Instruct, Apache-2.0)

Distilled: every label was blended with the soft labels (probabilities) of Cloudflare/clef-flash (Apache-2.0); the target is 50% original label + 50% teacher.

Held-out results (split of the training pool)

accuracy Brier ECE
zero-shot (before training) 0.670 0.409 0.182
trained 0.927 0.040 0.025
trained + calibrated 0.927 0.038 0.049

These are in-distribution numbers. Benchmark on your own data (and recalibrate on a few hundred labelled examples from it with systemone calibrate) before letting the model act automatically.

General benchmark (headline)

systemone bench --suite general: 1,328 held-out items, 30 questions in 7 families of state. Labels come from public held-out datasets, or from exact rules for generated records, logs and agent steps; no model labels them. The score is macro accuracy (the mean over questions). Both models ran on one RTX 3090, 5/5 identical on repeat.

Family kenning-large-v0.4 Cloudflare clef-flash
All 30 questions 0.625 0.813
Text: evidence, sentiment, toxicity, prompt injection, intent, topic (14) 0.754 0.841
Conversations (1) 0.958 1.000
Answer quality: helpful, correct (2) 0.377 0.447
Agent decisions: tool call matches, should call a function, task completed (3) 0.553 0.793
Records: rules over JSON with numbers and dates (7) 0.535 0.857
Tables: is a statement true (1) 0.480 0.860
Logs: is a service failing, which one (2) 0.290 0.740
Latency p50 34 ms 149 ms

Kenning is close to Clef on text, about 20 times smaller and 4 times faster. On structured state (records, tables, agent steps, logs) it is far behind and often near chance. If your state looks like that, measure it on your own data first, or train it for your problem. Per-task results and method: Kenning docs.

Email suites

Recorded with systemone bench on suites never used for training. Automated is the share of items decided without a person (p >= 0.9 or <= 0.1); a threat auto-closed is a positive item the model was sure was negative.

Suite Items Accuracy Brier ECE Automated Threats auto-closed Latency p50
Modern emails 2 (held out) 20 0.750 0.151 0.186 20% 0 47 ms
Modern emails 20 1.000 0.024 0.106 70% 0 49 ms
Phishing dataset 50 0.780 / 0.820 0.151 0.211 44% 0 88 ms
Layouts (4 state shapes) 200 0.780 / 0.760 / 0.760 / 0.760 0.151 0.211 38% 0 58 ms
Out of domain 180 0.917 / 0.583 / 0.867 0.086 0.172 30% 1 48 ms

Where a suite asks several questions, accuracy lists each (out of domain: spam / emotion / news topic; layouts: one per layout). Comparisons with Cloudflare Clef and TypeSafe Jev on the same suites are in the Kenning docs.

Use

pip install "systemone-client[local]"
from systemone import Kenning, Noul, Choice

model = Kenning.from_pretrained("systemonedev/kenning-large-v0.4")  # in-process, runs on CPU or GPU
answer = model.system_one(state={"ticket": "I was charged twice."},
                          questions={"billing": Noul("Is this about billing?")})
print(answer.nouls["billing"].noul)

Or serve it with SystemOne Builder's kenning service and call POST /v1/systemone with systemone.Client.

Limitations

  • Structured state is the weak spot. Rules over records, numbers and dates, tables, agent steps and logs score 0.29–0.55 on the general benchmark (Clef: 0.74–0.86). Its training data was mostly short text.
  • 512 tokens per (state, answer) pair. Longer states are truncated, so a decision that depends on the end of a long log or document is not seen.
  • Calibration was fitted on the training distribution; probabilities on other data are scores until recalibrated (systemone calibrate refits the temperatures on your labelled data).
  • Answer-quality judgements (is an answer helpful or correct) are hard for it and for Clef.
  • Determinism: identical requests give identical answers on the same hardware and software; results can differ across GPUs or library versions.

Licence

The weights are released under the Apache License 2.0 (LICENSE). Upstream licences of the base model and of every training source are listed in NOTICE.md; some are share-alike (CC-BY-SA-3.0), so keep NOTICE.md with the weights.

Kenning implements a wire format compatible with TypeSafe AI's System One API. It is not affiliated with or endorsed by TypeSafe AI, and was not trained on TypeSafe outputs.

Downloads last month
63
Safetensors
Model size
0.4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for systemonedev/kenning-large-v0.4

Finetuned
(3)
this model