Statim Decide EN Large

A decision model for Statim, the native C++ engine for typed decisions: ask any text a choice, a score or a yes/no question and get calibrated answers from one forward pass, on CPU or GPU, without Python at runtime. Version 0.5.0, fine-tuned from convaiinnovations/laya (ModernBERT-large encoder).

One support ticket, three typed answers, one forward pass: the 60-second film.

Quick start

# Statim release binary: https://github.com/BEKO2210/statim/releases
huggingface-cli download Beko2210/statim-decide-en-large statim-decide-en-large-q8_0.gguf --local-dir models
./statim serve -m english=models/statim-decide-en-large-q8_0.gguf --port 8080
curl -s localhost:8080/v1/systemone -d '{
  "state": {"subject": "Duplicate charge on invoice #4411",
            "body": "We were billed twice for March. Please refund the second charge."},
  "questions": {
    "department": {"type": "choice", "instructions": "Which department should handle this?",
      "criteria": {"billing": "invoices, refunds", "technical": "bugs, outages", "sales": "pricing, contracts"}},
    "urgency": {"type": "score", "instructions": "How urgent is this request?",
      "criteria": ["not urgent", "soon", "critical"]},
    "refund": {"type": "noul", "instructions": "Does the user explicitly request a refund?"}}}'

Open http://127.0.0.1:8080/ for the playground. API reference: docs/API.md.

Files

File Size Use
statim-decide-en-large-f32.gguf 1.58 GB reference precision, exact on GPU
statim-decide-en-large-q8_0.gguf 0.45 GB recommended for CPU: 4x smaller
checkpoint/ 0.85 GB Laya-format checkpoint for fine-tuning and the Python reference

Checksums in SHA256SUMS. The f32 file reproduces the reference implementation within 1e-4 on Statim's parity tests; q8_0 is smaller and faster on CPU with slightly different logits.

Evaluation

Measured by Statim's no-harm gate (tools/finetune/gate.py) on held-out test data the model selection never looked at. Against the checkpoint it was trained from, on 54 held-out suites: 11 significant gains, 43 within noise, 0 regressions (two combined binomial standard errors).

Suite Role This model Base checkpoint Protocol
typed-decisions trained 0.7680 0.3610 test split, first 2,000 decisions; its train split is replay data
Banking77 trained 0.9280 0.5500 test split, first 2,000 rows, all 77 intents in one question
MASSIVE intents (English) trained 0.8667 0.5333 150 seeded stratified English test rows
AG News held out 0.9390 0.9425 zero-shot (never trained on), first 2,000 test rows
DAIR Emotion held out 0.5880 0.5945 zero-shot, first 2,000 test rows
HWU64 intents held out 0.8333 0.6067 English, 150 rows; rows overlapping MASSIVE removed
SIB-200 topics held out 0.7133 0.7267 English, 150 rows; zero-shot
Sentiment held out 0.6733 0.6200 English, 150 rows; zero-shot
HateCheck held out 0.7800 0.7667 English, 150 rows; zero-shot
Belebele reading held out 0.4533 0.3800 English, 150 rows; zero-shot

Published systems under the same protocol, for orientation: typed-decisions meraGPT 0.768, laya-typed-decisions 0.766, Jev 0.727; AG News zero-shot Laya 0.950, GPT-3 (CARP) 0.926, Jev 0.881; Banking77 supervised MPNet 0.941. Sources: docs/ROADMAP.md.

All 54 held-out suites
Suite Accuracy Rows
amazon_massive_intent/ar 0.1067 150
amazon_massive_intent/de 0.4600 150
amazon_massive_intent/en 0.8667 150
amazon_massive_intent/es 0.5333 150
amazon_massive_intent/fr 0.5667 150
amazon_massive_intent/hi 0.0467 150
amazon_massive_intent/it 0.4067 150
amazon_massive_intent/ja 0.4133 150
amazon_massive_intent/pl 0.2400 150
amazon_massive_intent/ru 0.3867 150
amazon_massive_intent/tr 0.1867 150
amazon_massive_intent/zh-CN 0.5467 150
belebele/ar 0.2800 150
belebele/de 0.4467 150
belebele/en 0.4533 150
belebele/hi 0.2333 150
farstail/fa 0.3933 150
go_emotions/en 0.5267 150
hwu64/en 0.8333 150
indonli/id 0.5133 150
multi_hatecheck/ar 0.5000 150
multi_hatecheck/de 0.5133 150
multi_hatecheck/en 0.7800 150
multi_hatecheck/es 0.6267 150
multi_hatecheck/fr 0.5600 150
multi_hatecheck/hi 0.5400 150
multi_hatecheck/it 0.5533 150
multi_hatecheck/nl 0.5400 150
multi_hatecheck/pl 0.5133 150
multi_hatecheck/pt 0.5800 150
multi_hatecheck/zh 0.5933 150
multilingual_sentiments/ar 0.3667 150
multilingual_sentiments/de 0.4800 150
multilingual_sentiments/en 0.6733 150
multilingual_sentiments/es 0.5333 150
multilingual_sentiments/fr 0.5067 150
multilingual_sentiments/hi 0.4933 150
multilingual_sentiments/id 0.4733 150
multilingual_sentiments/it 0.3800 150
multilingual_sentiments/ja 0.4133 150
multilingual_sentiments/ms 0.4133 150
multilingual_sentiments/pt 0.4933 150
multilingual_sentiments/zh 0.6267 150
semrel/ar 0.2533 150
semrel/en 0.3133 150
semrel/hi 0.2333 150
sib200/ar 0.1867 150
sib200/de 0.6467 150
sib200/en 0.7133 150
sib200/hi 0.1733 150
test/ag_news 0.9390 2000
test/banking77 0.9280 2000
test/emotion 0.5880 2000
test/typed_decisions 0.7680 2000

Reproduce these numbers: REPRODUCE.md.

Training

Multi-task fine-tuning with train_multitask.py --clean from checkpoint laya, best epoch 11/raw selected on validation data only. Training data: only sources whose licence permits commercial use and imposes no ShareAlike or copyleft terms (Banking77, MASSIVE, typed-decisions replay, a licence-audited tasksource mixture, Nemotron-Safety-Guard, IndicGuard, MINDS-14, SNIPS), every one listed with its licence in DATA_LICENSES.md. Evaluation test rows were removed from the training data.

Intended use and limits

  • Classification-style decisions over short texts and JSON: routing, triage, moderation, intent, yes/no checks, ordinal ratings. It does not generate text.
  • Accuracy varies by task and language (see the table). Reading comprehension (Belebele) and semantic similarity are weak; do not use it for them without your own evaluation.
  • Use the confidence: Statim's min_confidence option marks low-confidence answers with escalate: true so a person can review them. Do not automate high-stakes decisions about people without human review.

Licence

The weights may be used under any one of: PolyForm Noncommercial 1.0.0, PolyForm Small Business 1.0.0 (free commercial use below 100 people and 1 M USD revenue), PolyForm Free Trial 1.0.0 (any company, fewer than 32 days), or a Statim commercial licence (COMMERCIAL.md). Texts in LICENSE-MODEL.md. The Statim engine is Apache-2.0.

Built on Laya (Apache-2.0) and ModernBERT-large (Apache-2.0). Training data attribution: Banking77 (Casanueva et al., 2020, PolyAI), MASSIVE (FitzGerald et al., 2022, Amazon), and the CC-BY sources in DATA_LICENSES.md. Statim is independent and not affiliated with the Laya authors.

Downloads last month
20
GGUF
Model size
0.4B params
Architecture
laya
Hardware compatibility
Log In to add your hardware

8-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Beko2210/statim-decide-en-large

Quantized
(42)
this model

Datasets used to train Beko2210/statim-decide-en-large

Evaluation results

  • accuracy on typed-decisions (test split, first 2,000 decisions; its train split is replay data)
    self-reported
    0.768
  • accuracy on Banking77 (test split, first 2,000 rows, all 77 intents in one question)
    self-reported
    0.928
  • accuracy on MASSIVE intents (English) (150 seeded stratified English test rows)
    self-reported
    0.867
  • accuracy on AG News (zero-shot (never trained on), first 2,000 test rows)
    self-reported
    0.939
  • accuracy on DAIR Emotion (zero-shot, first 2,000 test rows)
    self-reported
    0.588
  • accuracy on HWU64 intents (English, 150 rows; rows overlapping MASSIVE removed)
    self-reported
    0.833
  • accuracy on SIB-200 topics (English, 150 rows; zero-shot)
    self-reported
    0.713
  • accuracy on Sentiment (English, 150 rows; zero-shot)
    self-reported
    0.673