OpenDecider

OpenDecider-nano

Open, calibrated System 1 decision model. Give it a state (text, email, ticket or JSON) and typed questions (choice, score, noul); it returns a calibrated probability for every option in a single forward pass: 17 ms on an NVIDIA L40S, 18 ms on an Apple M4 Max, ~9 ms per question batched. ~400M parameters, Apache-2.0. It never generates text, so there is nothing to parse and nothing to hallucinate.

Beats Laya and TypeSafe Jev on typed-decisions: 0.796, against 0.766 for Laya's typed-decisions checkpoint (+0.030, 95% CI +0.014 to +0.044) and 0.754 for Jev, measured through TypeSafe's own API.

Installation

pip install opendecider

Python 3.10 or newer; Linux, Windows or macOS; CPU, NVIDIA (CUDA) or Apple Silicon (MPS). Platform notes are in the GitHub README.

Quickstart

from opendecider import load

model = load("manjunathshiva/opendecider-nano")   # 0.8 GB download on first use

state = "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
questions = {
    "department": {"type": "choice", "instructions": "Which department should handle this?",
                   "criteria": {"billing": "invoices, payments, refunds",
                                "technical": "bugs, outages, system errors",
                                "other": "everything else"}},
    "urgency": {"type": "score", "instructions": "How urgent is this?",
                "criteria": ["not urgent", "soon", "blocking"]},
    "churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"},
}

result = model.system_one(state, questions)
print(result["answers"]["department"]["choice"])   # billing        (probability 0.927)
print(result["answers"]["urgency"]["score"])       # 2 = blocking   (probability 0.604)
print(result["answers"]["churn_risk"]["noul"])     # 0.922 = probability the answer is yes

What's new in 0.1.0

  • First release of OpenDecider-nano and its 4B sibling, OpenDecider-small.
  • Jev measured directly through TypeSafe's own API on every benchmark, alongside Laya, CLM-8B and five frontier LLMs.
  • Identical results on Apple Silicon (MPS) and Linux + NVIDIA (CUDA).
  • Coming next (in development): OpenDecider-medium (Qwen3-30B-A3B) and OpenDecider-large (Qwen3-Next-80B-A3B), MLX builds, and a Colab notebook.

OpenDecider vs TypeSafe Jev, Laya, CLM-8B and frontier LLMs: typed-decisions, general decisions, Laya's battery, calibration, speed and open weights, same questions and same scorer

Highlighted: best in each column. typed-decisions scored with the Antz AI harness; OpenDecider-nano and Laya's typed-decisions checkpoint were fine-tuned on the train split, and the test split was never seen. Speeds: OpenDecider on an NVIDIA L40S, Laya on Apple Silicon, APIs include the network. Every number: COMPARISON.md.

OpenDecider versus TypeSafe Jev, Laya, CLM-8B and frontier LLMs

Will it fit?

Hardware Memory used Latency, one question Tested
Mac mini M4, 16 GB 2.0 GiB of the 11.8 GiB GPU budget 28 ms ✅
MacBook Pro M4 Max, 64 GB 2.0 GiB 18 ms ✅
NVIDIA L40S (Linux) ~2 GB 16 ms ✅
CPU only ~2 GB of RAM 0.1–0.7 s ✅ (the live demo runs on a basic CPU)

It should also fit any Mac with 8 GB and any NVIDIA GPU with 4 GB (not tested). Answers are identical across the tested machines to four decimals.

Architecture

  • Backbone: Ettin-encoder-400m (bidirectional, fully fine-tuned) + a small MLP decision head (Linear–GELU–LayerNorm–Linear).
  • Option markers: every option gets its own [MASK] token in question: …, [MASK] option 1, [MASK] option 2, …, input: <state>. The hidden state at each marker becomes one logit, softmaxed over that question's options. The answer space is defined at request time, so new schemas need no retraining.
  • No per-option token budget: options take the tokens they need and the state is truncated first (2,048 tokens in total), so a 78-option question costs one forward pass.
  • Batching: all questions in a call are answered in one padded batch.
  • Weights: stored in bf16 (0.8 GB), run in fp32 (2.0 GiB).

Training

Distillation from calibrated teachers. Two openly licensed teachers, Qwen3-235B-A22B-Instruct-2507 (Apache-2.0) and DeepSeek V4.1 Flash (MIT), scored every training question through token log-probabilities, each temperature-scaled on held-out gold labels before averaging; datasets with gold labels only use label-smoothed gold. Then a short fine-tune on the typed-decisions train split (100 train cases held out for model selection; the test split never used). No benchmark dataset below, or its family, is in the training data, and every training pool was checked for text overlap with all test sets (0 overlaps). No outputs of Claude or GPT models were used.

Benchmarks

Every model answered the same questions and was scored by the same code; TypeSafe Jev was measured through TypeSafe's own API. Full tables: COMPARISON.md.

Speed

questions per call NVIDIA L40S Apple M4 Max
1 16.1 ms 18.1 ms
5 24.4 ms (4.9 ms/q) 54.3 ms (10.9 ms/q)
10 42.9 ms (4.3 ms/q) 98.1 ms (9.8 ms/q)
50 189.5 ms (3.8 ms/q) 467 ms (9.3 ms/q)

TypeSafe Jev answered at a 404 ms median per question through its API in our runs.

OpenDecider-nano vs TypeSafe Jev and Laya

Benchmark / metric TypeSafe Jev 1.13 Laya Laya typed-decisions OpenDecider-nano
typed-decisions, 2,000 decisions 0.754 0.362 0.766 0.796
200 general decisions (BANKING77, BoolQ, Yelp, ChaosNLI) 0.730 0.545 0.570 0.680
Laya's application battery, 10 tasks 0.774 0.695 0.702 0.656
Laya's battery, the 5 tasks Laya was not trained on 0.803 0.579 0.609 0.656
Calibration error (ECE), general decisions 0.164 0.327 0.162 0.092
Median latency, 1 question 404 ms (API) 22 ms 21 ms 17 ms (L40S)
Weights closed API Apache-2.0 Apache-2.0 Apache-2.0

typed-decisions by question type

Scored with the Jev-vs-Laya harness published by Kameshwara Pavan kumar Mantha and the Antz AI team, joined question by question with their per-question results (0 gold-label mismatches).

model accuracy choice score yes/no vs Laya-td (95% CI)
OpenDecider-nano 0.796 0.762 0.769 0.867 +0.030 [+0.014, +0.044]
Laya typed-decisions 0.766 0.733 0.723 0.857 –
TypeSafe Jev 1.13 0.754 0.737 0.701 0.843 −0.012 [−0.034, +0.009]

KL divergence from the gold probability distributions: 0.079 (Jev 1.155).

Honest limits

  • Laya is better on the datasets it was trained on (spam 0.99, phishing 0.98, AG News 0.95), and leads Laya's battery overall (0.695 vs 0.656). Phishing (0.63) is this model's weakest task.
  • Jev leads Laya's application battery (0.774 vs 0.656; phishing 0.90, spam 0.985, routing 0.975).
  • Jev leads on general decisions (0.730 vs 0.680), especially BoolQ-style yes/no reading (0.94 vs 0.74). For the most accurate zero-shot decisions use OpenDecider-small (0.735).
  • English only so far; no multilingual evaluation has been run.
  • Descriptions help: very terse or cryptic option labels are harder; give options a short description when you can.

Links

Apache 2.0 · Base model Ettin-encoder-400m (MIT) · Manjunath Janardhan

Downloads last month
38
Safetensors
Model size
0.4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for manjunathshiva/opendecider-nano

Finetuned
(11)
this model

Datasets used to train manjunathshiva/opendecider-nano

Space using manjunathshiva/opendecider-nano 1

Collection including manjunathshiva/opendecider-nano

Article mentioning manjunathshiva/opendecider-nano