solvi-ai/solvi-large (preview)
A 396M-parameter ModernBERT-large cross-encoder that answers typed questions about a text or a JSON state for
solvi: one option, several options, ordered scores, yes / no, not stated, spans,
evidence quotes, rankings and numbers (bins), with a calibrated confidence and an act / escalate signal per answer.
Same format as the 150M solvi-base (solvi_decide v2, subformat l14g typed v2); this is the larger, stronger variant.
Honest summary. It is clearly better than our previous deciders on every test, and better than GLiNER2.5-Decide on typed questions over JSON states — but it does not beat GLiNER2.5-Decide on zero-shot choice questions (Fast Decisions dev 59.4% vs 62.9%). With 64 labelled examples per head it reaches 63.0%. It is also slower: 137 ms per question on a CPU (ONNX fp16).
Results (held-out tests; never used for training or checkpoint selection)
Criteria were registered before training (research log L19; 4 of 10 met for this model — the misses are under Limitations). The Fast Decisions numbers are on the public dev split (100 rows × 17 domains) in our harness; Fastino's official scores (GLiNER2.5-Decide 60.1, …) are on a hidden test split and are not directly comparable.
Choice questions — Fast Decisions (dev, macro over domains):
| model | params | zero-shot | + bias correction, unlabelled (Zc) | + 32 labelled (Sc) | + 64 labelled (Sc) |
|---|---|---|---|---|---|
| GLiNER2.5-Decide (fastino) | 340M | 62.9% | — | — | — |
| solvi-large (this model) | 396M | 59.4% | 60.5% | 62.0% | 63.0% |
| solvi-base (distilled student) | 150M | 56.3% | 58.2% | 60.2% | 60.4% |
| L14g (previous base, not published) | 150M | 56.2% | 57.4% | 58.8% | — |
| solvi L14d (previous clean decider) | 150M | 57.3% | 59.5% | 61.4% | 61.9% |
| knowledgator/gliformer-large-v1 (our run) | 576M | 43.9% | — | — | — |
| Laya (our earlier run) | — | 48.9% | — | — | — |
Single-label heads alone: 63.4%; exact multi-label sets: 26.3%.
Typed questions over JSON states — typed-decisions test (2000 questions):
| model | zero-shot | fitted per process, 32 / 300 examples |
|---|---|---|
| solvi-large | 54.5% | 63.1% / 64.9% |
| GLiNER2.5-Decide (our run) | 50.3% | — |
| L14g (previous base) | 45.1% | 58.8% / 61.3% |
| Laya base | 36.0% | fine-tuned on this dataset: 76.7% |
typed-decisions gold comes from a ~4B teacher's judgement, so "accuracy" here means agreement with that teacher.
New kinds on real text:
| test | measure | result |
|---|---|---|
| ContractNLI test (in-distribution: ContractNLI train was used for training) | yes / no / not stated | 88.8% |
| ContractNLI | the evidence quote supports the answer (entailed / contradicted pairs) | 62.0% (previous base L14g: 23%) |
| Taskmaster-2, held-out slots | span exact / token F1 | 70.2% / 0.83 |
| Taskmaster-2 | "not stated" (noisy gold) | 35.5% |
Synthetic states with 9 held-out schemas: all kinds 97.9% (prose rendering 98.6%); yes / no / not stated 98.9%; span exact 99.4%; evidence quotes literally correct 100% and supporting 99.5%; ranking NDCG@3 0.983; number median bin 95.8%.
Act / escalate signal (AUROC of "the answer is right"): ContractNLI 0.85, Fast Decisions 0.76, Taskmaster-2 0.74,
typed-decisions 0.67, synthetic 0.93. The shipped act calibrator is fitted on development data and is over-confident on new
real text — set the threshold on your data with act_guard (below).
Calibration (ECE): ContractNLI 0.026, Taskmaster-2 0.062, synthetic ≤ 0.02 per kind (rank 0.21), typed-decisions 0.21, Fast Decisions 0.21 (fit on your task before trusting zero-shot confidences).
Speed (CPU, 4 threads, idle machine): ONNX fp16, one question with 10 options and ~20 words: 137 ms (fp32 128 ms); a question over a long JSON state: 450–740 ms. ONNX = PyTorch on 300 of 300 Fast Decisions heads.
Several questions per pass (block layout) — answers agree with one-question-per-pass only 90.8% (typed-decisions) /
90.2% (Fast Decisions) / 97.5% (synthetic) of the time, so solvi_decide.json ships with multi_question.enabled = false.
The block ONNX export (onnx/model_block_fp16.onnx) is included for experiments (same answer as PyTorch 99.3%).
Training
- Initial weights: answerdotai/ModernBERT-large (Apache-2.0).
- Recipe: 200 minutes on one GPU per run (≈ 9.7k updates, effective batch 128, layer-wise LR decay 0.92), a first stage with 75% classification data, then the full mix; 35% of batches in the block layout with a block/full consistency loss. This model is the weight average of the best and the last checkpoint of that run, chosen by a development metric (8 permissive held-out classification sets + development splits of the training sources) — never by the tests above. Three sibling runs (a capacity-aware single-stage mix, a run without GLiNER distillation, a clean-data annealing of the second run) scored lower on that metric; averaging across runs made it worse.
- Teachers (labels only on the license-clean texts below; never on test data), all Apache-2.0: Qwen/Qwen2.5-32B-Instruct, Qwen/Qwen2.5-14B-Instruct, mistralai/Mistral-Small-24B-Instruct-2501, fastino/GLiNER2.5-Decide. Soft targets = vote shares; questions without ≥ 2 agreeing votes were dropped. On heads with real gold the consensus was right 87.5% of the time (GLiNER alone 73.3%).
- Data agent: Qwen2.5-32B-Instruct generated 22.7k texts with classification schemas over 76 business / content domains × 28 task types × 26 layouts; each head was verified by Qwen2.5-14B and Mistral-Small (kept at ≥ 2 of 3). The domain matrix was written from scratch and overlaps the public category descriptions of Fast Decisions; no Fast Decisions row was used, and all generated texts were de-duplicated against the Fast Decisions dev split (0 near-duplicates).
- De-duplication: exact + MinHash (Jaccard ≥ 0.5) against Fast Decisions dev, typed-decisions test, ContractNLI test and Taskmaster-2 test texts; 216 ContractNLI-train windows too close to test contracts were removed.
Training data and attribution
The full manifest with licenses, URLs and counts is LICENSES.json. Policy: permissive, public-domain and share-alike data (CC BY-SA,
with attribution here); excluded: non-commercial, research-only, unlicensed, or terms that forbid training.
| source | license | attribution |
|---|---|---|
| solvi synthetic corpora (operational texts, typed states, judgment states; own generators, no external text) | Apache-2.0 | solvi authors |
| SQuAD 2.0 (train) | CC BY-SA 4.0 | Pranav Rajpurkar, Robin Jia, Percy Liang: "Know What You Don't Know: Unanswerable Questions for SQuAD", ACL 2018 |
| BoolQ (train) | CC BY-SA 3.0 | Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, Kristina Toutanova: "BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions", NAACL 2019 |
| CFPB Consumer Complaint Database | CC0-1.0 | Consumer Financial Protection Bureau |
| English Wikinews (2023-07-28) | CC BY 2.5 | Wikinews contributors, https://en.wikinews.org |
| arXiv abstracts (metadata) | CC0-1.0 | arXiv.org |
| Taskmaster-1 | CC BY 4.0 | Google Research |
| MultiWOZ 2.2 | MIT | Budzianowski et al. |
| Amazon polarity | Apache-2.0 | Zhang et al.; McAuley & Leskovec |
| GoEmotions | Apache-2.0 | Google Research |
| LEDGAR, UNFAIR-ToS (LexGLUE) | CC BY 4.0 | Tuggener et al.; Lippi et al.; Chalkidis et al. |
| SMS Spam Collection | CC BY 4.0 | Almeida & Hidalgo (UCI) |
| Toxic conversations 50k (Civil Comments / Jigsaw) | CC BY 4.0 | Jigsaw / Conversation AI; MTEB |
| deepset prompt-injections | Apache-2.0 | deepset |
| CLINC150 | CC BY 3.0 | Larson et al. |
| Banking77 | CC BY 4.0 | PolyAI |
| MASSIVE (en-US) | CC BY 4.0 | Amazon Science |
| SNIPS NLU | Apache-2.0 | Snips |
| Bitext customer-support, retail-banking, travel, insurance | CDLA-Sharing-1.0 | Bitext Innovations |
| VitaminC (includes FEVER-based claims) | CC BY-SA 3.0 | Tal Schuster, Adam Fisch, Regina Barzilay: "Get Your Vitamin C! Robust Fact Verification with Contrastive Evidence", NAACL 2021; Wikipedia contributors | | ContractNLI (train split only) | CC BY 4.0 | Yuta Koreeda, Christopher D. Manning, "ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts", Findings of EMNLP 2021; Hitachi America, Ltd. | | LLM-generated texts and labels (Qwen2.5-32B / 14B-Instruct, Mistral-Small-24B-Instruct-2501) | Apache-2.0 (model licenses) | Qwen team (Alibaba Cloud); Mistral AI |
SQuAD 2.0, BoolQ and VitaminC are share-alike datasets: this card is their attribution. The model weights are released under Apache-2.0; if you redistribute the datasets themselves, their own licenses apply.
Checkpoint selection only (not trained on): DBpedia-14 (CC BY-SA 3.0), HuffPost News Category (CC BY 4.0), zeroshot twitter-financial-news topic / sentiment (MIT), Civil Comments (CC0), poem_sentiment (CC BY 4.0), LexGLUE SCOTUS (CC BY 4.0), Amazon counterfactual (CC BY 4.0).
Independent benchmarks (as released, zero-shot)
| benchmark | solvi-large | solvi-base | others (their published numbers or the benchmark's leaderboard) |
|---|---|---|---|
| jabr classifier-benchmark v2 (49 tasks, 866 cases), macro | 0.694 | 0.602 | Jev 0.966, GLiNER2.5-Decide 0.739, Von 0.720, GLiNER2 0.684, Laya 0.583 |
| decision-models-under-pressure: accuracy with 128 options | 45.1% | 50.8% | Jev 60.0%, best other open model 41.0%, Laya 38.5% |
| same: answers changed by reordering 64 options | 40.9% as listed · 0.5% in solvi ≥ 0.5.1 (sorted order by default) | 24.1% as listed | Jev 14.6%, Laya 49.4% |
| same: accuracy lost to near-duplicate distractors (64 options) | 0.260 | 0.247 | Jev 0.105, Laya 0.345 |
Yes / no questions are the weak spot on jabr: the model ranks them reasonably but says "yes" 36% of the time where the gold
labels have 50% — calibrate the threshold on your data (act_guard, fit).
Escalation with a guarantee (solvi ≥ 0.5.1)
The act threshold shipped in solvi_decide.json was fitted on development data and does not keep its promise on new real
text. Calibrate on a few hundred labelled examples of your own stream instead:
part.act_guard(examples, risk=0.10) # P(answered alone and wrong) ≤ 10% of all questions, for inputs like the examples
Measured with 300 calibration examples per data set and 200 random splits (the rest of the set is the test). "Answered" is the share decided without a person, "error" the error among those, "risk" the share of all questions answered alone and wrong — the number the guarantee is about:
| data set | shipped "10%" threshold: answered / error | act_guard(risk=0.10): answered / error / risk |
act_guard(risk=0.05) |
|---|---|---|---|
| typed-decisions | 78% / 39% | 32% / 30% / 9.8% | 18% / 27% / 5.0% |
| Taskmaster-2 | 81% / 32% | 50% / 20% / 10.0% | 34% / 14% / 4.9% |
| ContractNLI | 98% / 10% | 97% / 10% / 9.6% | 85% / 6% / 5.1% |
| JSON, 9 held-out schemas | 100% / 2% | 99.6% / 1.8% / 1.8% | 99.7% / 1.9% / 1.9% |
The guarantee holds on every set; how much can be automated depends on how hard the questions are. It holds for inputs like the calibration examples, not under a shift of domain: recalibrate when your inputs change.
Evaluation only — never trained on: Fast Decisions (fastino, dev split), typed-decisions test, ContractNLI test, Taskmaster-2 held-out domains.
Limitations
- Not better than GLiNER2.5-Decide on zero-shot choice questions (59.4% vs 62.9% on Fast Decisions dev). It reaches GLiNER's zero-shot level only after fitting on ~64 labelled examples per head.
- Slower than 100 ms on a CPU (137 ms per short question, ONNX fp16); use solvi-base (50 ms) where latency matters.
- The act / escalate calibrator is over-confident on new real text (at the "≤ 10% error" threshold fitted on development data the
real error was 32–39% on Fast Decisions, typed-decisions and Taskmaster-2; 10% on ContractNLI). Use
act_guardon your data (solvi ≥ 0.5.1): the risk it promises held on every set we measured. - Several questions per pass disagree with one-question-per-pass in ~9% of answers on real states — disabled by default.
- ContractNLI numbers are in-distribution (train split used); evidence on other document types is weaker.
- "Not stated" in dialogues is under-recognised (35.5% on Taskmaster-2); rank confidences are poorly calibrated (ECE 0.21).
- English only; inputs longer than 512 tokens (1024 in the block layout) are truncated — only the text, never the question.
- Not for medical, legal or credit decisions on its own; in solvi those rules belong in hard checks with escalation to a person.
Files
model.safetensors (bf16), onnx/model_fp16.onnx (full layout: input_ids, attention_mask → logits [B, L, 6]),
onnx/model_block_fp16.onnx (block layout: input_ids, position_ids, full_attention_mask, sliding_attention_mask),
onnx/parity.json, config.json, tokenizer.json, tokenizer_config.json, solvi_decide.json (capabilities, temperatures,
act calibrator, multi_question.enabled = false), l14g_format.py and l14f_format.py (the exact input / output format),
LICENSES.json (training data manifest).
Citation
@software{solvi,
title = {solvi: verifiable decision systems from catalogs of functions and checks},
author = {mxkuzn and solvi contributors},
year = {2026},
url = {https://github.com/solvi-ai/solvi}
}
- Downloads last month
- 26