verdict
A small typed decision model: give it a state (text or JSON) and a question with fixed options, get back a calibrated probability for every option. Nothing is generated. Three question types: Noul (yes/no), Choice (pick one of the options you define), Score (a level on a scale, plus its expected value).
Verdict follows the "System One" decision-model pattern popularized by TypeSafe AI's Jev, and its question types mirror Jev's. It is an independent project, not affiliated with or endorsed by TypeSafe AI.
These are the trainable weights only (11.3M parameters: LoRA r=16 adapters, the decision head and three fitted temperatures). They load on top of Qwen/Qwen3.5-0.8B-Base with the code in github.com/danieltrt/verdict.
Use
git clone https://github.com/danieltrt/verdict && cd verdict && uv sync
from verdict import Noul, Choice, Score, Verdict
vd = Verdict.load(checkpoint="drramos/verdict")
angry, team, urgency = vd.ask(
"Charged twice for March again. Refund it today or I'm cancelling.",
[Noul("Is the customer angry?"),
Choice("Which team should handle this?", ["billing", "technical support", "sales"]),
Score("How urgent is this ticket?", 1, 5)],
)
team.value, team.confidence, urgency.expected
Architecture
Interactive diagram, worked example and math: https://claude.ai/artifact/8ns3mSxQ6neq1y2Efagynq
Follows vllm-sr's Decision-1.0-Eos-0.8B:
the prompt is the state, the question type, the question, each option, then a fixed
suffix ending in Decision:. The hidden state at each option's last token (c_i) and
at the final token (g) go through a shared head, a 256-d bilinear term plus a small
GELU MLP, giving one logit per option. A per-type temperature and a softmax turn the
logits into probabilities.
Training
One RTX 5090, 33 minutes, 2 epochs. About 21k examples from 14 public datasets (BoolQ, SciQ, RACE, SQuAD 2.0 answerability, BANKING77, MASSIVE, AG News, MultiNLI, PAWS, Yelp, emotion, STS-B, ARC, CommonsenseQA), 4 synthetic families with exact labels (arithmetic as choice and yes/no, ticket urgency, policy rules with "cannot tell"), and 1,000 synthetic service logs of 1k–16k tokens. Temperatures were fitted on a separate calibration split.
Results
Held-out test splits, 200 questions per family, after calibration:
| Type | Accuracy | ECE |
|---|---|---|
| Choice | 84.0% | 0.018 |
| Noul | 87.5% | 0.029 |
| Score | 70.7% (mean error 0.37 levels) | 0.030 |
Two families were never trained on: time questions (75.0%) and team routing
(64.0%, overconfident). Per-family numbers are in metrics.json.
Limitations
- Weak at reading tone and implied intent in longer business text (e.g. it missed "refund it today or I'm cancelling" as a cancellation threat).
- Confuses who acts with who is mentioned ("Ana will send the budget to finance").
- Date arithmetic over weeks is close to a coin flip.
- The long-context family is a single easy task (find one ERROR line); it shows the pipeline handles 16k-token states, not general long-document reasoning.
License note
The base model is Apache 2.0. Some training datasets carry their own terms, including non-commercial ones (for example the Yelp review data), so check them before any commercial use.
Model tree for drramos/verdict
Base model
Qwen/Qwen3.5-0.8B-Base