OpenDecision-Large

📄 Technical report: OpenDecision: Fast Schema-Conditioned Decisions with Packed Candidate Encoding (PDF; arXiv version coming soon) · Code

OpenDecision-Large is a 433.9M-parameter encoder that makes decisions over options you define at call time: intent, routing, triage, sentiment, document type, yes/no gates and multi-label tags — several questions about the same text in one call. It does not generate text; it returns a probability for every option.

This is the accuracy-first model (cross-encoder: one encoder pass per option). For the lowest cost per decision use OpenDecision-Large-Packed.

Usage

pip install "opendecision @ git+https://github.com/tokz-labs/OpenDecision"
from opendecision import OpenDecision, Choice, Bool, MultiLabel

model = OpenDecision.from_pretrained("Tokz-labs/OpenDecision-Large")
out = model.decide(
    "I was charged twice for my May invoice and the app keeps logging me out.",
    {
        "queue": Choice(["billing", "technical", "account", "sales"]),
        "tags": MultiLabel(["double_charge", "login_issue", "refund_request", "outage"]),
        "urgent": Bool(),
    },
)
print(out["queue"].best, out["tags"].selected, out["urgent"].policy)

Benchmarks

Classification suite, complete test sets, macro-F1. Every system is an open model run through its own package on identical inputs, label names and questions (GLiNER2.5 with its best interface per dataset):

Model AG News CLINC150 IMDb Rotten Tomatoes XNLI Average
OpenDecision-Large 87.37 76.38 94.84 92.03 88.93* 87.91
OpenDecision-Large-Packed 86.70 75.95 94.59 89.96 88.97* 87.23
SemIf (Qwen3.5-4B, 4B) 87.28 78.22 95.27 88.46 82.36 86.32
JevK5 v0.3 (4B) 85.45 79.86 95.06 87.60 83.40 86.27
Laya (421M) 92.42 47.57 92.16 87.99 87.22 81.47
GLiFormer large-v1 (576M) 82.88 59.49 94.19 84.06 82.03 80.53
GLiNER2.5-Decide (340M) 71.37 66.19 90.02 85.83 47.65 72.21

* SNLI, MultiNLI and WANLI are in the training data: XNLI is not zero-shot for this model. Without XNLI it still has the highest average (87.66 vs 87.31 for SemIf). AG News, CLINC150, IMDb and Rotten Tomatoes contribute no training text, but guided our choice of training data. On one NVIDIA L4 it needs 77.4 GPU-seconds per 1,000 of these decisions, against 267.4 for JevK5 and 275.0 for SemIf.

Label sets that no training data targeted (macro-F1; GoEmotions: exact match). Here the 4B models and Laya generalize better:

Model TREC Emotion Subjectivity Average of 3 GoEmotions
JevK5 v0.3 (4B) 88.00 51.66 76.54 72.07 3.85
SemIf (Qwen3.5-4B, 4B) 83.36 47.99 61.53 64.29 0.87
Laya (421M) 83.00 48.54 54.79 62.11 10.26
OpenDecision-Large 65.13 49.15 56.71 57.00 10.83
OpenDecision-Large-Packed 62.95 48.30 55.68 55.64 13.30
GLiNER2.5-Decide (340M) 59.09 53.33 55.93 56.12 30.85
GLiFormer large-v1 (576M) 57.97 41.71 50.50 50.06 34.51

The final training stage adds 125,142 examples of 15,942 classification tasks invented by an LLM, with these task families excluded; it raised the three-task average from 45.97 to 57.00. On GoEmotions the model selects too many labels; a tuned threshold on the policy works better than the default outcome threshold. For a new kind of decision, fine-tune on a few hundred labeled examples (python -m opendecision.cli.train).

Technical details

  • Backbone: DeBERTa-v3-large; two linear heads on each option's [CLS] vector — a policy score (softmax over the option set) and a candidate-local outcome logit (sigmoid).
  • Input: one [CLS] question + text [SEP] option [SEP] sequence per option, up to 512 tokens; options truncated to 32 tokens.
  • Policy probabilities are overconfident; dividing policy logits by a temperature of about 2.8 fixes most of it (technical report, §3.4).
  • Throughput: 77.4 GPU-seconds per 1,000 decisions on the classification suite, one NVIDIA L4 in bf16.
  • Training: two stages from DeBERTa-v3-large on LLM-generated decisions and LLM-invented classification tasks (OpenAI models), the Fast Decisions public development split, MultiNLI, WANLI and samples of DBpedia-14, BANKING77, SNLI, WikiText-103 and 20 Newsgroups (Mitchell 1997, CC BY 4.0). Every source allows commercial use under its license; see the technical report.

Scope and limitations

English only. Not a chat, reasoning or question-answering model. Inputs beyond 512 tokens are truncated. The benchmarks above were also used to select training mixtures. Most training text was generated or labeled by OpenAI models.

Citation

@misc{alwarawreh2026opendecision,
  title  = {OpenDecision: Fast Schema-Conditioned Decisions with Packed Candidate Encoding},
  author = {Alwarawreh, Abdallah},
  year   = {2026},
  note   = {Tokz Labs technical report},
  url    = {https://github.com/tokz-labs/OpenDecision/releases/latest/download/OpenDecision.pdf}
}
Downloads last month
26
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Tokz-labs/OpenDecision-Large

Finetuned
(316)
this model

Datasets used to train Tokz-labs/OpenDecision-Large