Statim Decide EN Large
A decision model for Statim, the native C++ engine for typed decisions: ask any text a
choice, a score or a yes/no question and get calibrated answers from one forward pass, on
CPU or GPU, without Python at runtime. Version 0.5.0, fine-tuned from
convaiinnovations/laya (ModernBERT-large encoder).
One support ticket, three typed answers, one forward pass: the 60-second film.
Quick start
# Statim release binary: https://github.com/BEKO2210/statim/releases
huggingface-cli download Beko2210/statim-decide-en-large statim-decide-en-large-q8_0.gguf --local-dir models
./statim serve -m english=models/statim-decide-en-large-q8_0.gguf --port 8080
curl -s localhost:8080/v1/systemone -d '{
"state": {"subject": "Duplicate charge on invoice #4411",
"body": "We were billed twice for March. Please refund the second charge."},
"questions": {
"department": {"type": "choice", "instructions": "Which department should handle this?",
"criteria": {"billing": "invoices, refunds", "technical": "bugs, outages", "sales": "pricing, contracts"}},
"urgency": {"type": "score", "instructions": "How urgent is this request?",
"criteria": ["not urgent", "soon", "critical"]},
"refund": {"type": "noul", "instructions": "Does the user explicitly request a refund?"}}}'
Open http://127.0.0.1:8080/ for the playground. API reference: docs/API.md.
Files
| File | Size | Use |
|---|---|---|
statim-decide-en-large-f32.gguf |
1.58 GB | reference precision, exact on GPU |
statim-decide-en-large-q8_0.gguf |
0.45 GB | recommended for CPU: 4x smaller |
checkpoint/ |
0.85 GB | Laya-format checkpoint for fine-tuning and the Python reference |
Checksums in SHA256SUMS. The f32 file reproduces the reference implementation within 1e-4 on
Statim's parity tests; q8_0 is smaller and faster on CPU with slightly different logits.
Evaluation
Measured by Statim's no-harm gate (tools/finetune/gate.py)
on held-out test data the model selection never looked at. Against the checkpoint it was trained from, on 54 held-out suites: 11 significant gains, 43 within noise, 0 regressions (two combined binomial standard errors).
| Suite | Role | This model | Base checkpoint | Protocol |
|---|---|---|---|---|
| typed-decisions | trained | 0.7680 | 0.3610 | test split, first 2,000 decisions; its train split is replay data |
| Banking77 | trained | 0.9280 | 0.5500 | test split, first 2,000 rows, all 77 intents in one question |
| MASSIVE intents (English) | trained | 0.8667 | 0.5333 | 150 seeded stratified English test rows |
| AG News | held out | 0.9390 | 0.9425 | zero-shot (never trained on), first 2,000 test rows |
| DAIR Emotion | held out | 0.5880 | 0.5945 | zero-shot, first 2,000 test rows |
| HWU64 intents | held out | 0.8333 | 0.6067 | English, 150 rows; rows overlapping MASSIVE removed |
| SIB-200 topics | held out | 0.7133 | 0.7267 | English, 150 rows; zero-shot |
| Sentiment | held out | 0.6733 | 0.6200 | English, 150 rows; zero-shot |
| HateCheck | held out | 0.7800 | 0.7667 | English, 150 rows; zero-shot |
| Belebele reading | held out | 0.4533 | 0.3800 | English, 150 rows; zero-shot |
Published systems under the same protocol, for orientation: typed-decisions meraGPT 0.768, laya-typed-decisions 0.766, Jev 0.727; AG News zero-shot Laya 0.950, GPT-3 (CARP) 0.926, Jev 0.881; Banking77 supervised MPNet 0.941. Sources: docs/ROADMAP.md.
All 54 held-out suites
| Suite | Accuracy | Rows |
|---|---|---|
amazon_massive_intent/ar |
0.1067 | 150 |
amazon_massive_intent/de |
0.4600 | 150 |
amazon_massive_intent/en |
0.8667 | 150 |
amazon_massive_intent/es |
0.5333 | 150 |
amazon_massive_intent/fr |
0.5667 | 150 |
amazon_massive_intent/hi |
0.0467 | 150 |
amazon_massive_intent/it |
0.4067 | 150 |
amazon_massive_intent/ja |
0.4133 | 150 |
amazon_massive_intent/pl |
0.2400 | 150 |
amazon_massive_intent/ru |
0.3867 | 150 |
amazon_massive_intent/tr |
0.1867 | 150 |
amazon_massive_intent/zh-CN |
0.5467 | 150 |
belebele/ar |
0.2800 | 150 |
belebele/de |
0.4467 | 150 |
belebele/en |
0.4533 | 150 |
belebele/hi |
0.2333 | 150 |
farstail/fa |
0.3933 | 150 |
go_emotions/en |
0.5267 | 150 |
hwu64/en |
0.8333 | 150 |
indonli/id |
0.5133 | 150 |
multi_hatecheck/ar |
0.5000 | 150 |
multi_hatecheck/de |
0.5133 | 150 |
multi_hatecheck/en |
0.7800 | 150 |
multi_hatecheck/es |
0.6267 | 150 |
multi_hatecheck/fr |
0.5600 | 150 |
multi_hatecheck/hi |
0.5400 | 150 |
multi_hatecheck/it |
0.5533 | 150 |
multi_hatecheck/nl |
0.5400 | 150 |
multi_hatecheck/pl |
0.5133 | 150 |
multi_hatecheck/pt |
0.5800 | 150 |
multi_hatecheck/zh |
0.5933 | 150 |
multilingual_sentiments/ar |
0.3667 | 150 |
multilingual_sentiments/de |
0.4800 | 150 |
multilingual_sentiments/en |
0.6733 | 150 |
multilingual_sentiments/es |
0.5333 | 150 |
multilingual_sentiments/fr |
0.5067 | 150 |
multilingual_sentiments/hi |
0.4933 | 150 |
multilingual_sentiments/id |
0.4733 | 150 |
multilingual_sentiments/it |
0.3800 | 150 |
multilingual_sentiments/ja |
0.4133 | 150 |
multilingual_sentiments/ms |
0.4133 | 150 |
multilingual_sentiments/pt |
0.4933 | 150 |
multilingual_sentiments/zh |
0.6267 | 150 |
semrel/ar |
0.2533 | 150 |
semrel/en |
0.3133 | 150 |
semrel/hi |
0.2333 | 150 |
sib200/ar |
0.1867 | 150 |
sib200/de |
0.6467 | 150 |
sib200/en |
0.7133 | 150 |
sib200/hi |
0.1733 | 150 |
test/ag_news |
0.9390 | 2000 |
test/banking77 |
0.9280 | 2000 |
test/emotion |
0.5880 | 2000 |
test/typed_decisions |
0.7680 | 2000 |
Reproduce these numbers: REPRODUCE.md.
Training
Multi-task fine-tuning with train_multitask.py --clean from checkpoint laya, best
epoch 11/raw selected on validation data only. Training data: only sources whose
licence permits commercial use and imposes no ShareAlike or copyleft terms (Banking77, MASSIVE,
typed-decisions replay, a licence-audited tasksource mixture, Nemotron-Safety-Guard, IndicGuard,
MINDS-14, SNIPS), every one listed with its licence in
DATA_LICENSES.md. Evaluation test rows were removed from the training data.
Intended use and limits
- Classification-style decisions over short texts and JSON: routing, triage, moderation, intent, yes/no checks, ordinal ratings. It does not generate text.
- Accuracy varies by task and language (see the table). Reading comprehension (Belebele) and semantic similarity are weak; do not use it for them without your own evaluation.
- Use the confidence: Statim's
min_confidenceoption marks low-confidence answers withescalate: trueso a person can review them. Do not automate high-stakes decisions about people without human review.
Licence
The weights may be used under any one of: PolyForm Noncommercial 1.0.0, PolyForm Small Business 1.0.0 (free commercial use below 100 people and 1 M USD revenue), PolyForm Free Trial 1.0.0 (any company, fewer than 32 days), or a Statim commercial licence (COMMERCIAL.md). Texts in LICENSE-MODEL.md. The Statim engine is Apache-2.0.
Built on Laya (Apache-2.0) and ModernBERT-large (Apache-2.0). Training data attribution: Banking77 (Casanueva et al., 2020, PolyAI), MASSIVE (FitzGerald et al., 2022, Amazon), and the CC-BY sources in DATA_LICENSES.md. Statim is independent and not affiliated with the Laya authors.
- Downloads last month
- 20
8-bit
32-bit
Model tree for Beko2210/statim-decide-en-large
Base model
convaiinnovations/layaDatasets used to train Beko2210/statim-decide-en-large
AmazonScience/massive
PolyAI/minds14
Evaluation results
- accuracy on typed-decisions (test split, first 2,000 decisions; its train split is replay data)self-reported0.768
- accuracy on Banking77 (test split, first 2,000 rows, all 77 intents in one question)self-reported0.928
- accuracy on MASSIVE intents (English) (150 seeded stratified English test rows)self-reported0.867
- accuracy on AG News (zero-shot (never trained on), first 2,000 test rows)self-reported0.939
- accuracy on DAIR Emotion (zero-shot, first 2,000 test rows)self-reported0.588
- accuracy on HWU64 intents (English, 150 rows; rows overlapping MASSIVE removed)self-reported0.833
- accuracy on SIB-200 topics (English, 150 rows; zero-shot)self-reported0.713
- accuracy on Sentiment (English, 150 rows; zero-shot)self-reported0.673