OpenDecider

OpenDecider-small

Open, calibrated System 1 decision model for decisions it has never seen. Give it a state (text, email, ticket or JSON) and typed questions (choice, score, noul); it returns a calibrated probability for every option: 40 ms on an NVIDIA L40S, and it runs on a 16 GB Mac mini (tested). A 4B LoRA adapter on Qwen3-4B-Instruct-2507, Apache-2.0.

Zero-shot, it beats TypeSafe Jev and Laya on general decisions (0.735 vs 0.730 and 0.545) and ties Laya's best checkpoint on Laya's own application battery (0.702 vs 0.702), winning the five tasks Laya was not trained on by 13–16 points. Best-calibrated model you can run yourself (ECE 0.087; Jev 0.164, Laya 0.327).

Installation

pip install "opendecider[small]"

Python 3.10 or newer; Linux, Windows or macOS; NVIDIA (CUDA) or Apple Silicon (MPS) recommended. Downloads Qwen3-4B-Instruct-2507 (8 GB) plus this adapter (126 MB) on first use. Platform notes are in the GitHub README.

Quickstart

from opendecider import load, Choice, Noul

model = load("manjunathshiva/opendecider-small")

r = model.system_one(
    {"message": "I took out cash abroad and the exchange rate is wrong."},
    {"intent": Choice("Which banking intent is this?",
                      ["wrong_exchange_rate_for_cash_withdrawal", "card_payment_fee_charged",
                       "cash_withdrawal_charge", "declined_cash_withdrawal"]),
     "complaint": Noul("Is the customer complaining?")})
print(r["answers"]["intent"]["choice"], r["answers"]["intent"]["probabilities"])
print(r["answers"]["complaint"]["noul"])   # probability the answer is yes

What's new in 0.1.0

  • First release of OpenDecider-small and its fast ~400M sibling, OpenDecider-nano.
  • Jev measured directly through TypeSafe's own API on every benchmark, alongside Laya, CLM-8B and five frontier LLMs.
  • Identical results on Apple Silicon (MPS) and Linux + NVIDIA (CUDA).
  • Coming next (in development): OpenDecider-medium (Qwen3-30B-A3B) and OpenDecider-large (Qwen3-Next-80B-A3B), MLX builds, and a Colab notebook.

OpenDecider vs TypeSafe Jev, Laya, CLM-8B and frontier LLMs: typed-decisions, general decisions, Laya's battery, calibration, speed and open weights, same questions and same scorer

Highlighted: best in each column. typed-decisions scored with the Antz AI harness; OpenDecider-nano and Laya's typed-decisions checkpoint were fine-tuned on the train split, and the test split was never seen. Speeds: OpenDecider on an NVIDIA L40S, Laya on Apple Silicon, APIs include the network. Every number: COMPARISON.md.

OpenDecider versus TypeSafe Jev, Laya, CLM-8B and frontier LLMs

Will it fit?

Hardware Memory used Latency, one question Tested
Mac mini M4, 16 GB 8.9 GiB of the 11.8 GiB GPU budget 280 ms ✅
MacBook Pro M4 Max, 64 GB 8.9 GiB 141 ms ✅
NVIDIA L40S (Linux) ~9 GB (bf16) 38 ms ✅
CPU only (fp32) ~17 GB of RAM slow not recommended

A 16 GB Mac is enough (tested on an M4 Mac mini with ~3 GiB to spare). NVIDIA: a GPU with 12 GB or more. Answers are identical across these machines to four decimals.

Architecture

  • Backbone: Qwen3-4B-Instruct-2507 with a LoRA adapter (r = 16, alpha 32, all linear projections), merged at load.
  • Reading a decision: options are lettered, and one forward pass gives the probability of each letter as the next token. Above 26 options, each option name's log-probability after the shared prompt.
  • No generation: nothing to parse, and every answer is a full probability distribution.

Training

Distillation from calibrated teachers. Two openly licensed teachers, Qwen3-235B-A22B-Instruct-2507 (Apache-2.0) and DeepSeek V4.1 Flash (MIT), scored every training question through token log-probabilities, each temperature-scaled on held-out gold labels before averaging; datasets with gold labels only use label-smoothed gold. This model never saw typed-decisions or any other benchmark dataset below (or its family), and every training pool was checked for text overlap with all test sets (0 overlaps). No outputs of Claude or GPT models were used.

Benchmarks

Every model answered the same questions and was scored by the same code; TypeSafe Jev was measured through TypeSafe's own API. Full tables: COMPARISON.md.

Speed

questions per call NVIDIA L40S Apple M4 Max
1 37.6 ms 141 ms
5 190.1 ms (38.0 ms/q) 680 ms (136 ms/q)
10 388.2 ms (38.8 ms/q) 1.37 s (137 ms/q)
50 1.94 s (38.7 ms/q) 6.86 s (137 ms/q)

Memory: 8.9 GiB (bf16), tested on a 16 GB Mac mini (M4). TypeSafe Jev: 404 ms median per question through its API.

OpenDecider-small vs TypeSafe Jev and Laya (zero-shot)

Benchmark / metric TypeSafe Jev 1.13 Laya Laya typed-decisions OpenDecider-small
200 general decisions (BANKING77, BoolQ, Yelp, ChaosNLI) 0.730 0.545 0.570 0.735
Laya's application battery, 10 tasks 0.774 0.695 0.702 0.702
Laya's battery, the 5 tasks Laya was not trained on 0.803 0.579 0.609 0.743
BANKING77, 77 labels (Laya's battery) 0.845 0.425 0.492 0.748
typed-decisions, 2,000 decisions 0.754 0.362 0.766 (fine-tuned) 0.672 (zero-shot)
Calibration error (ECE), general decisions 0.164 0.327 0.162 0.087
Distance from the human label spread (ChaosNLI JSD) 0.148 0.174 0.111 0.040
Median latency, 1 question 404 ms (API) 22 ms 21 ms 40 ms (L40S)

Against frontier LLMs (same 200 general decisions)

Model accuracy ECE median latency $ / 1,000 decisions
Claude Fable 5.1 0.840 0.064 4.27 s $11.81
GPT-6 Astra 0.790 0.119 2.22 s $6.96
DeepSeek V4.1 Flash 0.760 0.138 4.08 s $0.158
OpenDecider-small 0.735 0.087 40 ms self-hosted
TypeSafe Jev 1.13 0.730 0.164 404 ms $0.025
Qwen3-4B-Instruct-2507, untrained (this model's base) 0.700 0.289 – –

Distillation moved the base model from 0.700 to 0.735 and cut its calibration error from 0.289 to 0.087.

Honest limits

  • Laya is better on the datasets it was trained on (AG News, spam, phishing, support triage). Phishing (0.63) is this model's weakest task.
  • Jev leads Laya's application battery (0.774 vs 0.702, and 0.803 vs 0.743 on the tasks Laya was not trained on), with phishing 0.90, spam 0.985 and routing 0.975.
  • Jev leads on typed-decisions zero-shot (0.754 vs 0.672); the fine-tuned OpenDecider-nano passes both (0.796).
  • Frontier LLMs are more accurate (0.745–0.84), at 25–160× the latency and a per-call bill.
  • Answers questions one at a time; use OpenDecider-nano when you need many decisions per second.
  • English only so far; no multilingual evaluation has been run.

Links

Apache 2.0 · Base model Qwen3-4B-Instruct-2507 (Apache-2.0) · Manjunath Janardhan

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for manjunathshiva/opendecider-small

Adapter
(5721)
this model
Quantizations
2 models

Datasets used to train manjunathshiva/opendecider-small

Collection including manjunathshiva/opendecider-small

Article mentioning manjunathshiva/opendecider-small