OpenDecider

OpenDecider-medium-td

The most accurate OpenDecider, and the closest of all tested systems to human judgement. A 30B mixture-of-experts decision model (Qwen3-30B-A3B-Instruct-2507, 3B active, with a LoRA adapter): ask typed questions (choice, score, noul) about any text or JSON and get a calibrated probability for every option, with no text generation to parse. Apache-2.0, for NVIDIA GPUs.

  • 200 general decisions none of these models trained on: 0.765, ahead of TypeSafe Jev (0.730) and every other model you can run yourself; only Claude Fable 5.1 (0.840) and GPT-6 Astra (0.790) score higher.
  • Closest to the human label spread on ChaosNLI (100 human votes per item): JSD 0.035, against 0.148 for Jev.
  • typed-decisions: 0.788, against 0.766 for Laya's typed-decisions checkpoint (+0.022, 95% CI +0.005 to +0.040) and 0.754 for Jev.

GitHub PyPI version Collection Full comparison License

Installation

pip install torch --index-url https://download.pytorch.org/whl/cu128
pip install "opendecider[small]>=0.1.2"

Version 0.1.2 or newer is needed: it spreads the model across all visible GPUs.

Quickstart

from opendecider import load

model = load("manjunathshiva/opendecider-medium-td")   # downloads the 61 GB base on first use
r = model.system_one(
    {"invoice_id": "INV-2291", "vendor": "Acme Supplies", "amount": 4820.00, "currency": "USD",
     "po_number": None, "due": "2026-09-15", "note": "Second reminder, now 12 days overdue."},
    {"action": {"type": "choice", "instructions": "What should accounts payable do with this invoice?",
                "criteria": {"approve": "pay it", "hold": "hold for a missing purchase order", "reject": "not a valid invoice"}},
     "risk": {"type": "score", "instructions": "How risky is paying this invoice?",
              "criteria": ["low", "medium", "high"]},
     "needs_review": {"type": "noul", "instructions": "Should a human review this before payment?"}})
for name, a in r["answers"].items():
    print(name, a["probabilities"])

When to use which model

model best for
OpenDecider-medium-td (this) the highest accuracy on decisions it has never seen, and probabilities closest to how people disagree; NVIDIA, ~61 GB of GPU memory
OpenDecider-small the best calibration, on a 16 GB Mac or one GPU
OpenDecider-small-td business workflows (triage, invoices, security alerts, agent traces) on a 16 GB Mac or one GPU
OpenDecider-nano speed: 16 ms per question, ~400M parameters, runs on CPU

Benchmarks

Every model answered the same questions and was scored by the same code (benchmark harness). TypeSafe Jev was measured through TypeSafe's own API.

benchmark TypeSafe Jev 1.13 Laya typed-decisions OpenDecider-nano OpenDecider-small OpenDecider-small-td OpenDecider-medium-td
typed-decisions (2,000 decisions) 0.754 0.766 0.796 0.671 0.792 0.788
200 general decisions 0.730 0.570 0.680 0.735 0.715 0.765
Laya's application battery (10 tasks) 0.774 0.702 0.656 0.702 0.703 0.725
calibration error (ECE) ↓ 0.164 0.162 0.092 0.087 0.107 0.110
distance from human votes (ChaosNLI JSD) ↓ 0.148 0.111 0.045 0.040 0.040 0.035
median latency, 1 question 404 ms (API) 21 ms 16 ms (L40S) 40 ms (L40S) 40 ms (L40S) 214 ms (4× L40S)

typed-decisions scored with the Jev-vs-Laya harness published by Kameshwara Pavan kumar Mantha and the Antz AI team. Laya's typed-decisions checkpoint, nano, small-td and medium-td were fine-tuned on the train split; the test split was never used.

Against frontier LLMs (same 200 general decisions)

model accuracy ECE ↓ median latency
Claude Fable 5.1 0.840 0.064 4.27 s
GPT-6 Astra 0.790 0.119 2.22 s
OpenDecider-medium-td 0.765 0.110 214 ms
DeepSeek V4.1 Flash 0.760 0.138 4.08 s
MiniMax M3 0.755 0.112 1.02 s
Qwen3-30B-A3B-Instruct-2507, untrained (this model's base) 0.745 0.233 –
TypeSafe Jev 1.13 0.730 0.164 404 ms

Training moved the base model from 0.745 to 0.765 and cut its calibration error from 0.233 to 0.110. The 200-item set is about ±3 points, so medium-td, DeepSeek V4.1 Flash and MiniMax M3 are close.

Automating only the confident decisions

benchmark model all decisions most confident 70% most confident 50%
typed-decisions OpenDecider-medium-td 0.788 0.896 0.948
typed-decisions TypeSafe Jev 1.13 0.754 0.839 0.882
general (200) OpenDecider-medium-td 0.765 0.807 0.820
general (200) TypeSafe Jev 1.13 0.730 0.829 0.860

Where Jev leads

  • Laya's application battery (0.774 vs 0.725): phishing (0.897 vs 0.652), jailbreak detection (0.940 vs 0.762), spam (0.985 vs 0.950), model routing (0.975 vs 0.935) and 77-label BANKING77 (0.845 vs 0.785).
  • Ranking its own confidence on general decisions (0.860 vs 0.820 on the most confident half).

medium-td leads Jev on toxicity moderation (0.802 vs 0.665), general decisions, typed-decisions, calibration and agreement with human votes.

Will it fit?

61 GB of bf16 weights, spread automatically across all visible NVIDIA GPUs. Tested on 4× NVIDIA L40S (48 GB each); any set of GPUs with about 64 GB or more in total should work. No Mac build: a 4-bit MLX version of the medium model (before the typed-decisions fine-tune) scored 0.725 on general decisions, no better than OpenDecider-small-mlx-8bit (0.730), which needs 4.5 GB. On a Mac, use that instead.

Training

  1. Distillation from two calibrated, openly licensed teachers, Qwen3-235B-A22B-Instruct-2507 (Apache-2.0) and DeepSeek V4.1 Flash (MIT), on ~190K decision questions; each teacher was temperature-scaled on held-out gold labels before averaging.
  2. A short fine-tune (700 steps) on the typed-decisions train split mixed 1:1 with general data, with 100 train cases held out for model selection.

LoRA r = 16, α = 32 on the attention projections (q, k, v, o); the experts are frozen. Trained on AWS SageMaker (4× NVIDIA L40S). No benchmark dataset or its family is in the training data (0 text overlaps with any test set), and no outputs of Claude or GPT models were used.

Limitations

  • Phishing is the weakest task (0.65 on Laya's battery vs Jev's 0.90).
  • Answers one question per forward pass, 214 ms each on 4× L40S. Use nano when you need many decisions per second.
  • English only so far, and single training seeds.

Links

Apache 2.0 · Base model Qwen3-30B-A3B-Instruct-2507 (Apache-2.0) · Manjunath Janardhan

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for manjunathshiva/opendecider-medium-td

Adapter
(127)
this model

Datasets used to train manjunathshiva/opendecider-medium-td

Collection including manjunathshiva/opendecider-medium-td