Instructions to use manjunathshiva/opendecider-small with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use manjunathshiva/opendecider-small with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B-Instruct-2507") model = PeftModel.from_pretrained(base_model, "manjunathshiva/opendecider-small") - Notebooks
- Google Colab
- Kaggle
OpenDecider-small
Open, calibrated System 1 decision model for decisions it has never seen. Give it a state (text, email, ticket
or JSON) and typed questions (choice, score, noul); it returns a calibrated probability for every option:
40 ms on an NVIDIA L40S, and it runs on a 16 GB Mac mini (tested). A 4B LoRA adapter on Qwen3-4B-Instruct-2507, Apache-2.0.
Zero-shot, it beats TypeSafe Jev and Laya on general decisions (0.735 vs 0.730 and 0.545) and ties Laya's best checkpoint on Laya's own application battery (0.702 vs 0.702), winning the five tasks Laya was not trained on by 13–16 points. Best-calibrated model you can run yourself (ECE 0.087; Jev 0.164, Laya 0.327).
Installation
pip install "opendecider[small]"
Python 3.10 or newer; Linux, Windows or macOS; NVIDIA (CUDA) or Apple Silicon (MPS) recommended. Downloads Qwen3-4B-Instruct-2507 (8 GB) plus this adapter (126 MB) on first use. Platform notes are in the GitHub README.
Quickstart
from opendecider import load, Choice, Noul
model = load("manjunathshiva/opendecider-small")
r = model.system_one(
{"message": "I took out cash abroad and the exchange rate is wrong."},
{"intent": Choice("Which banking intent is this?",
["wrong_exchange_rate_for_cash_withdrawal", "card_payment_fee_charged",
"cash_withdrawal_charge", "declined_cash_withdrawal"]),
"complaint": Noul("Is the customer complaining?")})
print(r["answers"]["intent"]["choice"], r["answers"]["intent"]["probabilities"])
print(r["answers"]["complaint"]["noul"]) # probability the answer is yes
What's new in 0.1.0
- First release of OpenDecider-small and its fast ~400M sibling, OpenDecider-nano.
- Jev measured directly through TypeSafe's own API on every benchmark, alongside Laya, CLM-8B and five frontier LLMs.
- Identical results on Apple Silicon (MPS) and Linux + NVIDIA (CUDA).
- Coming next (in development): OpenDecider-medium (Qwen3-30B-A3B) and OpenDecider-large (Qwen3-Next-80B-A3B), MLX builds, and a Colab notebook.
Highlighted: best in each column. typed-decisions scored with the Antz AI harness; OpenDecider-nano and Laya's typed-decisions checkpoint were fine-tuned on the train split, and the test split was never seen. Speeds: OpenDecider on an NVIDIA L40S, Laya on Apple Silicon, APIs include the network. Every number: COMPARISON.md.
Will it fit?
| Hardware | Memory used | Latency, one question | Tested |
|---|---|---|---|
| Mac mini M4, 16 GB | 8.9 GiB of the 11.8 GiB GPU budget | 280 ms | ✅ |
| MacBook Pro M4 Max, 64 GB | 8.9 GiB | 141 ms | ✅ |
| NVIDIA L40S (Linux) | ~9 GB (bf16) | 38 ms | ✅ |
| CPU only (fp32) | ~17 GB of RAM | slow | not recommended |
A 16 GB Mac is enough (tested on an M4 Mac mini with ~3 GiB to spare). NVIDIA: a GPU with 12 GB or more. Answers are identical across these machines to four decimals.
Architecture
- Backbone: Qwen3-4B-Instruct-2507 with a LoRA adapter (r = 16, alpha 32, all linear projections), merged at load.
- Reading a decision: options are lettered, and one forward pass gives the probability of each letter as the next token. Above 26 options, each option name's log-probability after the shared prompt.
- No generation: nothing to parse, and every answer is a full probability distribution.
Training
Distillation from calibrated teachers. Two openly licensed teachers, Qwen3-235B-A22B-Instruct-2507 (Apache-2.0) and DeepSeek V4.1 Flash (MIT), scored every training question through token log-probabilities, each temperature-scaled on held-out gold labels before averaging; datasets with gold labels only use label-smoothed gold. This model never saw typed-decisions or any other benchmark dataset below (or its family), and every training pool was checked for text overlap with all test sets (0 overlaps). No outputs of Claude or GPT models were used.
Benchmarks
Every model answered the same questions and was scored by the same code; TypeSafe Jev was measured through TypeSafe's own API. Full tables: COMPARISON.md.
Speed
| questions per call | NVIDIA L40S | Apple M4 Max |
|---|---|---|
| 1 | 37.6 ms | 141 ms |
| 5 | 190.1 ms (38.0 ms/q) | 680 ms (136 ms/q) |
| 10 | 388.2 ms (38.8 ms/q) | 1.37 s (137 ms/q) |
| 50 | 1.94 s (38.7 ms/q) | 6.86 s (137 ms/q) |
Memory: 8.9 GiB (bf16), tested on a 16 GB Mac mini (M4). TypeSafe Jev: 404 ms median per question through its API.
OpenDecider-small vs TypeSafe Jev and Laya (zero-shot)
| Benchmark / metric | TypeSafe Jev 1.13 | Laya | Laya typed-decisions | OpenDecider-small |
|---|---|---|---|---|
| 200 general decisions (BANKING77, BoolQ, Yelp, ChaosNLI) | 0.730 | 0.545 | 0.570 | 0.735 |
| Laya's application battery, 10 tasks | 0.774 | 0.695 | 0.702 | 0.702 |
| Laya's battery, the 5 tasks Laya was not trained on | 0.803 | 0.579 | 0.609 | 0.743 |
| BANKING77, 77 labels (Laya's battery) | 0.845 | 0.425 | 0.492 | 0.748 |
| typed-decisions, 2,000 decisions | 0.754 | 0.362 | 0.766 (fine-tuned) | 0.672 (zero-shot) |
| Calibration error (ECE), general decisions | 0.164 | 0.327 | 0.162 | 0.087 |
| Distance from the human label spread (ChaosNLI JSD) | 0.148 | 0.174 | 0.111 | 0.040 |
| Median latency, 1 question | 404 ms (API) | 22 ms | 21 ms | 40 ms (L40S) |
Against frontier LLMs (same 200 general decisions)
| Model | accuracy | ECE | median latency | $ / 1,000 decisions |
|---|---|---|---|---|
| Claude Fable 5.1 | 0.840 | 0.064 | 4.27 s | $11.81 |
| GPT-6 Astra | 0.790 | 0.119 | 2.22 s | $6.96 |
| DeepSeek V4.1 Flash | 0.760 | 0.138 | 4.08 s | $0.158 |
| OpenDecider-small | 0.735 | 0.087 | 40 ms | self-hosted |
| TypeSafe Jev 1.13 | 0.730 | 0.164 | 404 ms | $0.025 |
| Qwen3-4B-Instruct-2507, untrained (this model's base) | 0.700 | 0.289 | – | – |
Distillation moved the base model from 0.700 to 0.735 and cut its calibration error from 0.289 to 0.087.
Honest limits
- Laya is better on the datasets it was trained on (AG News, spam, phishing, support triage). Phishing (0.63) is this model's weakest task.
- Jev leads Laya's application battery (0.774 vs 0.702, and 0.803 vs 0.743 on the tasks Laya was not trained on), with phishing 0.90, spam 0.985 and routing 0.975.
- Jev leads on typed-decisions zero-shot (0.754 vs 0.672); the fine-tuned OpenDecider-nano passes both (0.796).
- Frontier LLMs are more accurate (0.745–0.84), at 25–160× the latency and a per-call bill.
- Answers questions one at a time; use OpenDecider-nano when you need many decisions per second.
- English only so far; no multilingual evaluation has been run.
Links
- Live demo: https://huggingface.co/spaces/manjunathshiva/opendecider-demo
- GitHub: https://github.com/manjunathshiva/opendecider
- PyPI: https://pypi.org/project/opendecider/
- OpenDecider-nano (~400M): https://huggingface.co/manjunathshiva/opendecider-nano
- Collection: https://huggingface.co/collections/manjunathshiva/opendecider-6ab8c838909092518d50a9ea
- Training-data attributions: NOTICE
Apache 2.0 · Base model Qwen3-4B-Instruct-2507 (Apache-2.0) · Manjunath Janardhan
- Downloads last month
- -