verdict

A small typed decision model: give it a state (text or JSON) and a question with fixed options, get back a calibrated probability for every option. Nothing is generated. Three question types: Noul (yes/no), Choice (pick one of the options you define), Score (a level on a scale, plus its expected value).

Verdict follows the "System One" decision-model pattern popularized by TypeSafe AI's Jev, and its question types mirror Jev's. It is an independent project, not affiliated with or endorsed by TypeSafe AI.

These are the trainable weights only (11.3M parameters: LoRA r=16 adapters, the decision head and three fitted temperatures). They load on top of Qwen/Qwen3.5-0.8B-Base with the code in github.com/danieltrt/verdict.

Use

git clone https://github.com/danieltrt/verdict && cd verdict && uv sync
from verdict import Noul, Choice, Score, Verdict

vd = Verdict.load(checkpoint="drramos/verdict")
angry, team, urgency = vd.ask(
    "Charged twice for March again. Refund it today or I'm cancelling.",
    [Noul("Is the customer angry?"),
     Choice("Which team should handle this?", ["billing", "technical support", "sales"]),
     Score("How urgent is this ticket?", 1, 5)],
)
team.value, team.confidence, urgency.expected

Architecture

Interactive diagram, worked example and math: https://claude.ai/artifact/8ns3mSxQ6neq1y2Efagynq

Follows vllm-sr's Decision-1.0-Eos-0.8B: the prompt is the state, the question type, the question, each option, then a fixed suffix ending in Decision:. The hidden state at each option's last token (c_i) and at the final token (g) go through a shared head, a 256-d bilinear term plus a small GELU MLP, giving one logit per option. A per-type temperature and a softmax turn the logits into probabilities.

Training

One RTX 5090, 33 minutes, 2 epochs. About 21k examples from 14 public datasets (BoolQ, SciQ, RACE, SQuAD 2.0 answerability, BANKING77, MASSIVE, AG News, MultiNLI, PAWS, Yelp, emotion, STS-B, ARC, CommonsenseQA), 4 synthetic families with exact labels (arithmetic as choice and yes/no, ticket urgency, policy rules with "cannot tell"), and 1,000 synthetic service logs of 1k–16k tokens. Temperatures were fitted on a separate calibration split.

Results

Held-out test splits, 200 questions per family, after calibration:

Type Accuracy ECE
Choice 84.0% 0.018
Noul 87.5% 0.029
Score 70.7% (mean error 0.37 levels) 0.030

Two families were never trained on: time questions (75.0%) and team routing (64.0%, overconfident). Per-family numbers are in metrics.json.

Limitations

  • Weak at reading tone and implied intent in longer business text (e.g. it missed "refund it today or I'm cancelling" as a cancellation threat).
  • Confuses who acts with who is mentioned ("Ana will send the budget to finance").
  • Date arithmetic over weeks is close to a coin flip.
  • The long-context family is a single easy task (find one ERROR line); it shows the pipeline handles 16k-token states, not general long-document reasoning.

License note

The base model is Apache 2.0. Some training datasets carry their own terms, including non-commercial ones (for example the Yelp review data), so check them before any commercial use.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
11.3M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for drramos/verdict

Adapter
(31)
this model