DecidaBERT-large
DecidaBERT-large reads a state and a set of typed questions and answers all of them in one forward pass, with no generated tokens. Every answer is a probability distribution, not a sentence. It is a 421M parameter ModernBERT encoder that runs on a laptop GPU in tens of milliseconds, and it is built to be served by Decida.
The three question types:
| type | you give | you get back |
|---|---|---|
choice |
2 to 255 options, each with a description | a probability for every option |
score |
an ordered list of levels, lowest first | a probability for every level and the expected level |
noul |
a yes/no question about the state | the probability that the statement is true |
Quick start
uv tool install git+https://github.com/heldernoid/decida
decida serve --model decidabert=helmo/DecidaBERT-large --max-len 2048 # trained at 512, supported up to 2,048
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
"model": "decidabert",
"state": "Customer: I was billed twice for March and want my money back.",
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges, invoices, refunds",
"technical": "Bugs and outages",
"sales": "Pricing, demos, new seats"}},
"urgent": {"type": "noul", "instructions": "Is this urgent?"}
}}'
The response holds answers.team.probabilities, answers.urgent.noul and the latency. Many independent requests can be sent in one call to POST /v1/systemone/batch.
The weights use the same layout as Laya: the instructions and all option descriptions are packed after the state, and one marker per option is read out. They do not load with plain AutoModel, because the scoring head is custom. Use Decida, or read the loader in the Decida repository (src/decida/model/enc.py).
Results
The typed-decisions test set
Held-out test split of LocalLLaMA/typed-decisions: 400 cases, 2,000 questions. A prediction is right when its top answer equals the gold label. Intervals are 95% Wilson intervals. This model was trained on the train split of the same benchmark, so this is specialist mode.
| question type | questions | DecidaBERT-large | 95% interval | random guess | Qwen3-0.6B, zero-shot |
|---|---|---|---|---|---|
| choice | 600 | 0.722 | 0.685 to 0.756 | 0.233 | 0.350 |
| score | 800 | 0.694 | 0.661 to 0.725 | 0.244 | 0.281 |
yes/no (noul) |
600 | 0.830 | 0.798 to 0.858 | 0.500 | 0.550 |
| all | 2,000 | 0.743 | 0.723 to 0.762 | 0.383 |
For score questions the model is within one level of the gold level 95.5% of the time.
Calibration. Expected calibration error (10 bins, on the top answer) is 0.130 overall: 0.128 for choice, 0.137 for score and 0.122 for yes/no. The model is under-confident: its mean confidence (0.59 on choice, 0.56 on score, 0.71 on yes/no) is below its accuracy. No temperature scaling is applied (rl_agent_config.json records temperature 1.0). Qwen3-0.6B, read zero-shot through option letters, is the opposite: mean confidence 0.92 on choice at 0.35 accuracy.
How to read these numbers
- The gold labels are the average of three samples from a teacher model, so they are noisy. An earlier check put the teacher's agreement with itself at about 0.735, which means scores near 0.75 are close to the ceiling the labels allow.
- We make no claim that this model beats another one. For context only: on the same file and with the same scoring we measured Laya's typed-decisions checkpoint at 0.761 overall (choice 0.738, score 0.715, yes/no 0.847), served the same way as this model, which is higher than this model. The dataset author published 0.727 accuracy for TypeSafe Jev 1.13.0 on this test set, in generalist mode and with an aggregation we cannot confirm, so it is not a like-for-like number.
- Reproduce with
eval_typed_decisions.pyin this repository. It needs a running Decida server and the test file.
In the Decida testbench
Single runs, meant as behaviour examples and not as benchmarks. All of them show the same thing: the model does well when the state describes the situation in words, and badly when it is given raw numbers.
| task | result |
|---|---|
| Candies sorter: sort each candy into one of five bowls, 256 questions per request | 96.5% of candies in the right bowl when each candy is described in words, 19.2% when given raw RGB numbers |
| Tetris: pick the landing spot for each piece, every spot described in a sentence (3 games of 40 pieces) | 13.3 lines per game; its top choice matched an exact oracle's 83% of the time, and the oracle's answer was in its top three 99% of the time |
| Flappy: bird height put into words, 30 seconds | 22 pillars passed with no crash; with the height given as numbers, 13 crashes |
| Wikispeedia: click through real articles to a goal page (20 pairs, at most 20 clicks) | reached the goal in 13 of 20, where random clicking reaches it in 1 of 100; a larger hosted model reached all 20, so world knowledge is the limit here |
Speed on an Apple-silicon GPU: about 25 ms per single-question request (median, one request at a time), and about 2 seconds for 50 three-question requests sent in one batch call.
Writing questions that work
These come from what the model did well and badly in the testbench:
- Put the state in words, never in raw numbers. "The bird is far below the gap" works; "y=212, gap=140" does not.
- Ask what the text says, not what to do. "Which team does this ticket belong to?" beats "What should we do?".
- Describe every option. The option text is what the model matches the state against. A bare label is much weaker than a short description.
- Keep option lists short and only offer reachable options. It works up to 255 options, but 5 to 15 well described options is where it is strongest.
- Let your code do the measuring and choosing. Ask a perception question ("how far is the bird from the gap?") and apply your own threshold, rather than asking the model to decide everything.
- Watch the token budget. The model was trained on sequences of up to 512 tokens and runs up to 2,048. If you use long option descriptions, serve with
--head-tokens 1400 --option-tokens 160and checkx_decida.truncationin the response to see whether anything was cut.
Model details
| Architecture | ModernBERT-large encoder (28 layers, hidden size 1024) with a per-option scoring head, 421M parameters, stored in float32 (1.7 GB) |
| Base model | tasksource/ModernBERT-large-nli (Apache-2.0), itself a fine-tune of ModernBERT-large |
| Context | trained at 512 tokens, supported up to 2,048 |
| Language | English. Rows in non-Latin scripts were removed from the synthetic training data, so other scripts are untested |
| Input | one state, up to 256 questions per request |
Training
- Objective: supervised fine-tuning for 1,500 steps on target distributions (one hard label per synthetic record, soft distributions for typed-decisions), with cross-entropy plus 0.5 times a Brier term to keep the distributions calibrated.
- Data: 11,079 records, expanded to 15,879 training examples by using every question of each case.
- 9,879 synthetic records (3,282 choice, 3,304 score, 3,293 yes/no, one question each), published under the MIT licence as helmo/synthetic-typed-decisions. They cover 207 topic domains and were written by deepseek-ai/DeepSeek-V4-Flash-0731 through OpenRouter. Each output was parsed with a strict schema and rejected if a label was invalid, and rows in non-Latin scripts were dropped. The labels are written by that model, not by people.
- 1,200 training cases (five questions each) from LocalLLaMA/typed-decisions (Apache-2.0).
- Held out: the 400 test cases of LocalLLaMA/typed-decisions were never used for training, calibration or choosing a checkpoint.
- Hosted decision models: we did not call TypeSafe Jev, or use its outputs, to build the synthetic data or to train this model. The benchmark's gold labels come from its authors' teacher model, as described on the dataset card.
Limitations
- Little world knowledge. It matches the state against the options you give it. It does not know which Wikipedia page mentions another, or what a colour phrase like "tomato" looks like unless the options say so.
- Labels are noisy and model-written, both in the synthetic training data and in the benchmark's gold answers. Treat small differences between models on this benchmark with care.
- Specialist, not generalist. It was tuned on the typed-decisions distribution. Results on unrelated tasks can be much lower, especially with terse or unlabelled options.
- Under-confident on the benchmark. If you need probabilities to be calibrated on your own data, fit a temperature on a few hundred labelled examples of your own.
- Not for high-stakes decisions on its own. Use the probabilities, set a threshold that reflects the cost of being wrong, and send low-confidence cases to a person or a stronger model.
Files
| file | purpose |
|---|---|
model.safetensors |
the weights |
encoder/config.json |
the ModernBERT configuration |
tokenizer.json, tokenizer_config.json |
the tokenizer |
rl_agent_config.json |
calibration record (temperature 1.0, no scaling applied) |
eval_typed_decisions.py |
reproduces the results table |
License
Apache-2.0. The base model (tasksource/ModernBERT-large-nli), the benchmark and the source ModernBERT release are also Apache-2.0. See LICENSE and NOTICE.
Credits
Built on ModernBERT (Warner et al., Answer.AI and LightOn), the multi-task NLI fine-tune by tasksource (Sileo), and the typed-decisions benchmark by LocalLLaMA. The input layout follows Laya (Nandakishor M, Convai Innovations, Apache-2.0). Decida is not affiliated with TypeSafe AI, and the name Jev is used only to describe a published comparison.
Model tree for helmo/DecidaBERT-large
Base model
answerdotai/ModernBERT-large