OpenDecider-nano
Open, calibrated System 1 decision model. Give it a state (text, email, ticket or JSON) and typed questions
(choice, score, noul); it returns a calibrated probability for every option in a single forward pass:
17 ms on an NVIDIA L40S, 18 ms on an Apple M4 Max, ~9 ms per question batched. ~400M parameters, Apache-2.0.
It never generates text, so there is nothing to parse and nothing to hallucinate.
Beats Laya and TypeSafe Jev on typed-decisions: 0.796, against 0.766 for Laya's typed-decisions checkpoint (+0.030, 95% CI +0.014 to +0.044) and 0.754 for Jev, measured through TypeSafe's own API.
Installation
pip install opendecider
Python 3.10 or newer; Linux, Windows or macOS; CPU, NVIDIA (CUDA) or Apple Silicon (MPS). Platform notes are in the GitHub README.
Quickstart
from opendecider import load
model = load("manjunathshiva/opendecider-nano") # 0.8 GB download on first use
state = "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
questions = {
"department": {"type": "choice", "instructions": "Which department should handle this?",
"criteria": {"billing": "invoices, payments, refunds",
"technical": "bugs, outages, system errors",
"other": "everything else"}},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["not urgent", "soon", "blocking"]},
"churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"},
}
result = model.system_one(state, questions)
print(result["answers"]["department"]["choice"]) # billing (probability 0.927)
print(result["answers"]["urgency"]["score"]) # 2 = blocking (probability 0.604)
print(result["answers"]["churn_risk"]["noul"]) # 0.922 = probability the answer is yes
What's new in 0.1.0
- First release of OpenDecider-nano and its 4B sibling, OpenDecider-small.
- Jev measured directly through TypeSafe's own API on every benchmark, alongside Laya, CLM-8B and five frontier LLMs.
- Identical results on Apple Silicon (MPS) and Linux + NVIDIA (CUDA).
- Coming next (in development): OpenDecider-medium (Qwen3-30B-A3B) and OpenDecider-large (Qwen3-Next-80B-A3B), MLX builds, and a Colab notebook.
Highlighted: best in each column. typed-decisions scored with the Antz AI harness; OpenDecider-nano and Laya's typed-decisions checkpoint were fine-tuned on the train split, and the test split was never seen. Speeds: OpenDecider on an NVIDIA L40S, Laya on Apple Silicon, APIs include the network. Every number: COMPARISON.md.
Will it fit?
| Hardware | Memory used | Latency, one question | Tested |
|---|---|---|---|
| Mac mini M4, 16 GB | 2.0 GiB of the 11.8 GiB GPU budget | 28 ms | ✅ |
| MacBook Pro M4 Max, 64 GB | 2.0 GiB | 18 ms | ✅ |
| NVIDIA L40S (Linux) | ~2 GB | 16 ms | ✅ |
| CPU only | ~2 GB of RAM | 0.1–0.7 s | ✅ (the live demo runs on a basic CPU) |
It should also fit any Mac with 8 GB and any NVIDIA GPU with 4 GB (not tested). Answers are identical across the tested machines to four decimals.
Architecture
- Backbone: Ettin-encoder-400m (bidirectional, fully fine-tuned) + a small MLP decision head (Linear–GELU–LayerNorm–Linear).
- Option markers: every option gets its own
[MASK]token inquestion: …, [MASK] option 1, [MASK] option 2, …, input: <state>. The hidden state at each marker becomes one logit, softmaxed over that question's options. The answer space is defined at request time, so new schemas need no retraining. - No per-option token budget: options take the tokens they need and the state is truncated first (2,048 tokens in total), so a 78-option question costs one forward pass.
- Batching: all questions in a call are answered in one padded batch.
- Weights: stored in bf16 (0.8 GB), run in fp32 (2.0 GiB).
Training
Distillation from calibrated teachers. Two openly licensed teachers, Qwen3-235B-A22B-Instruct-2507 (Apache-2.0) and DeepSeek V4.1 Flash (MIT), scored every training question through token log-probabilities, each temperature-scaled on held-out gold labels before averaging; datasets with gold labels only use label-smoothed gold. Then a short fine-tune on the typed-decisions train split (100 train cases held out for model selection; the test split never used). No benchmark dataset below, or its family, is in the training data, and every training pool was checked for text overlap with all test sets (0 overlaps). No outputs of Claude or GPT models were used.
Benchmarks
Every model answered the same questions and was scored by the same code; TypeSafe Jev was measured through TypeSafe's own API. Full tables: COMPARISON.md.
Speed
| questions per call | NVIDIA L40S | Apple M4 Max |
|---|---|---|
| 1 | 16.1 ms | 18.1 ms |
| 5 | 24.4 ms (4.9 ms/q) | 54.3 ms (10.9 ms/q) |
| 10 | 42.9 ms (4.3 ms/q) | 98.1 ms (9.8 ms/q) |
| 50 | 189.5 ms (3.8 ms/q) | 467 ms (9.3 ms/q) |
TypeSafe Jev answered at a 404 ms median per question through its API in our runs.
OpenDecider-nano vs TypeSafe Jev and Laya
| Benchmark / metric | TypeSafe Jev 1.13 | Laya | Laya typed-decisions | OpenDecider-nano |
|---|---|---|---|---|
| typed-decisions, 2,000 decisions | 0.754 | 0.362 | 0.766 | 0.796 |
| 200 general decisions (BANKING77, BoolQ, Yelp, ChaosNLI) | 0.730 | 0.545 | 0.570 | 0.680 |
| Laya's application battery, 10 tasks | 0.774 | 0.695 | 0.702 | 0.656 |
| Laya's battery, the 5 tasks Laya was not trained on | 0.803 | 0.579 | 0.609 | 0.656 |
| Calibration error (ECE), general decisions | 0.164 | 0.327 | 0.162 | 0.092 |
| Median latency, 1 question | 404 ms (API) | 22 ms | 21 ms | 17 ms (L40S) |
| Weights | closed API | Apache-2.0 | Apache-2.0 | Apache-2.0 |
typed-decisions by question type
Scored with the Jev-vs-Laya harness published by Kameshwara Pavan kumar Mantha and the Antz AI team, joined question by question with their per-question results (0 gold-label mismatches).
| model | accuracy | choice | score | yes/no | vs Laya-td (95% CI) |
|---|---|---|---|---|---|
| OpenDecider-nano | 0.796 | 0.762 | 0.769 | 0.867 | +0.030 [+0.014, +0.044] |
| Laya typed-decisions | 0.766 | 0.733 | 0.723 | 0.857 | – |
| TypeSafe Jev 1.13 | 0.754 | 0.737 | 0.701 | 0.843 | −0.012 [−0.034, +0.009] |
KL divergence from the gold probability distributions: 0.079 (Jev 1.155).
Honest limits
- Laya is better on the datasets it was trained on (spam 0.99, phishing 0.98, AG News 0.95), and leads Laya's battery overall (0.695 vs 0.656). Phishing (0.63) is this model's weakest task.
- Jev leads Laya's application battery (0.774 vs 0.656; phishing 0.90, spam 0.985, routing 0.975).
- Jev leads on general decisions (0.730 vs 0.680), especially BoolQ-style yes/no reading (0.94 vs 0.74). For the most accurate zero-shot decisions use OpenDecider-small (0.735).
- English only so far; no multilingual evaluation has been run.
- Descriptions help: very terse or cryptic option labels are harder; give options a short description when you can.
Links
- Live demo: https://huggingface.co/spaces/manjunathshiva/opendecider-demo
- GitHub: https://github.com/manjunathshiva/opendecider
- PyPI: https://pypi.org/project/opendecider/
- OpenDecider-small (4B): https://huggingface.co/manjunathshiva/opendecider-small
- Collection: https://huggingface.co/collections/manjunathshiva/opendecider-6ab8c838909092518d50a9ea
- Training-data attributions: NOTICE
Apache 2.0 · Base model Ettin-encoder-400m (MIT) · Manjunath Janardhan
- Downloads last month
- 38
Model tree for manjunathshiva/opendecider-nano
Base model
jhu-clsp/ettin-encoder-400m