Instructions to use manjunathshiva/opendecider-medium-td with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use manjunathshiva/opendecider-medium-td with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-30B-A3B-Instruct-2507") model = PeftModel.from_pretrained(base_model, "manjunathshiva/opendecider-medium-td") - Notebooks
- Google Colab
- Kaggle
OpenDecider-medium-td
The most accurate OpenDecider, and the closest of all tested systems to human judgement. A 30B mixture-of-experts
decision model (Qwen3-30B-A3B-Instruct-2507, 3B active, with a LoRA adapter): ask typed questions (choice, score,
noul) about any text or JSON and get a calibrated probability for every option, with no text generation to parse.
Apache-2.0, for NVIDIA GPUs.
- 200 general decisions none of these models trained on: 0.765, ahead of TypeSafe Jev (0.730) and every other model you can run yourself; only Claude Fable 5.1 (0.840) and GPT-6 Astra (0.790) score higher.
- Closest to the human label spread on ChaosNLI (100 human votes per item): JSD 0.035, against 0.148 for Jev.
- typed-decisions: 0.788, against 0.766 for Laya's typed-decisions checkpoint (+0.022, 95% CI +0.005 to +0.040) and 0.754 for Jev.
Installation
pip install torch --index-url https://download.pytorch.org/whl/cu128
pip install "opendecider[small]>=0.1.2"
Version 0.1.2 or newer is needed: it spreads the model across all visible GPUs.
Quickstart
from opendecider import load
model = load("manjunathshiva/opendecider-medium-td") # downloads the 61 GB base on first use
r = model.system_one(
{"invoice_id": "INV-2291", "vendor": "Acme Supplies", "amount": 4820.00, "currency": "USD",
"po_number": None, "due": "2026-09-15", "note": "Second reminder, now 12 days overdue."},
{"action": {"type": "choice", "instructions": "What should accounts payable do with this invoice?",
"criteria": {"approve": "pay it", "hold": "hold for a missing purchase order", "reject": "not a valid invoice"}},
"risk": {"type": "score", "instructions": "How risky is paying this invoice?",
"criteria": ["low", "medium", "high"]},
"needs_review": {"type": "noul", "instructions": "Should a human review this before payment?"}})
for name, a in r["answers"].items():
print(name, a["probabilities"])
When to use which model
| model | best for |
|---|---|
| OpenDecider-medium-td (this) | the highest accuracy on decisions it has never seen, and probabilities closest to how people disagree; NVIDIA, ~61 GB of GPU memory |
| OpenDecider-small | the best calibration, on a 16 GB Mac or one GPU |
| OpenDecider-small-td | business workflows (triage, invoices, security alerts, agent traces) on a 16 GB Mac or one GPU |
| OpenDecider-nano | speed: 16 ms per question, ~400M parameters, runs on CPU |
Benchmarks
Every model answered the same questions and was scored by the same code (benchmark harness). TypeSafe Jev was measured through TypeSafe's own API.
| benchmark | TypeSafe Jev 1.13 | Laya typed-decisions | OpenDecider-nano | OpenDecider-small | OpenDecider-small-td | OpenDecider-medium-td |
|---|---|---|---|---|---|---|
| typed-decisions (2,000 decisions) | 0.754 | 0.766 | 0.796 | 0.671 | 0.792 | 0.788 |
| 200 general decisions | 0.730 | 0.570 | 0.680 | 0.735 | 0.715 | 0.765 |
| Laya's application battery (10 tasks) | 0.774 | 0.702 | 0.656 | 0.702 | 0.703 | 0.725 |
| calibration error (ECE) ↓ | 0.164 | 0.162 | 0.092 | 0.087 | 0.107 | 0.110 |
| distance from human votes (ChaosNLI JSD) ↓ | 0.148 | 0.111 | 0.045 | 0.040 | 0.040 | 0.035 |
| median latency, 1 question | 404 ms (API) | 21 ms | 16 ms (L40S) | 40 ms (L40S) | 40 ms (L40S) | 214 ms (4× L40S) |
typed-decisions scored with the Jev-vs-Laya harness published by Kameshwara Pavan kumar Mantha and the Antz AI team. Laya's typed-decisions checkpoint, nano, small-td and medium-td were fine-tuned on the train split; the test split was never used.
Against frontier LLMs (same 200 general decisions)
| model | accuracy | ECE ↓ | median latency |
|---|---|---|---|
| Claude Fable 5.1 | 0.840 | 0.064 | 4.27 s |
| GPT-6 Astra | 0.790 | 0.119 | 2.22 s |
| OpenDecider-medium-td | 0.765 | 0.110 | 214 ms |
| DeepSeek V4.1 Flash | 0.760 | 0.138 | 4.08 s |
| MiniMax M3 | 0.755 | 0.112 | 1.02 s |
| Qwen3-30B-A3B-Instruct-2507, untrained (this model's base) | 0.745 | 0.233 | – |
| TypeSafe Jev 1.13 | 0.730 | 0.164 | 404 ms |
Training moved the base model from 0.745 to 0.765 and cut its calibration error from 0.233 to 0.110. The 200-item set is about ±3 points, so medium-td, DeepSeek V4.1 Flash and MiniMax M3 are close.
Automating only the confident decisions
| benchmark | model | all decisions | most confident 70% | most confident 50% |
|---|---|---|---|---|
| typed-decisions | OpenDecider-medium-td | 0.788 | 0.896 | 0.948 |
| typed-decisions | TypeSafe Jev 1.13 | 0.754 | 0.839 | 0.882 |
| general (200) | OpenDecider-medium-td | 0.765 | 0.807 | 0.820 |
| general (200) | TypeSafe Jev 1.13 | 0.730 | 0.829 | 0.860 |
Where Jev leads
- Laya's application battery (0.774 vs 0.725): phishing (0.897 vs 0.652), jailbreak detection (0.940 vs 0.762), spam (0.985 vs 0.950), model routing (0.975 vs 0.935) and 77-label BANKING77 (0.845 vs 0.785).
- Ranking its own confidence on general decisions (0.860 vs 0.820 on the most confident half).
medium-td leads Jev on toxicity moderation (0.802 vs 0.665), general decisions, typed-decisions, calibration and agreement with human votes.
Will it fit?
61 GB of bf16 weights, spread automatically across all visible NVIDIA GPUs. Tested on 4× NVIDIA L40S (48 GB each); any set of GPUs with about 64 GB or more in total should work. No Mac build: a 4-bit MLX version of the medium model (before the typed-decisions fine-tune) scored 0.725 on general decisions, no better than OpenDecider-small-mlx-8bit (0.730), which needs 4.5 GB. On a Mac, use that instead.
Training
- Distillation from two calibrated, openly licensed teachers, Qwen3-235B-A22B-Instruct-2507 (Apache-2.0) and DeepSeek V4.1 Flash (MIT), on ~190K decision questions; each teacher was temperature-scaled on held-out gold labels before averaging.
- A short fine-tune (700 steps) on the typed-decisions train split mixed 1:1 with general data, with 100 train cases held out for model selection.
LoRA r = 16, α = 32 on the attention projections (q, k, v, o); the experts are frozen. Trained on AWS SageMaker (4× NVIDIA L40S). No benchmark dataset or its family is in the training data (0 text overlaps with any test set), and no outputs of Claude or GPT models were used.
Limitations
- Phishing is the weakest task (0.65 on Laya's battery vs Jev's 0.90).
- Answers one question per forward pass, 214 ms each on 4× L40S. Use nano when you need many decisions per second.
- English only so far, and single training seeds.
Links
- GitHub: https://github.com/manjunathshiva/opendecider
- Full comparison: https://github.com/manjunathshiva/opendecider/blob/main/COMPARISON.md
- Collection: https://huggingface.co/collections/manjunathshiva/opendecider-6ab8c838909092518d50a9ea
- Training-data attributions: NOTICE
Apache 2.0 · Base model Qwen3-30B-A3B-Instruct-2507 (Apache-2.0) · Manjunath Janardhan
- Downloads last month
- 10
Model tree for manjunathshiva/opendecider-medium-td
Base model
Qwen/Qwen3-30B-A3B-Instruct-2507