solvi-ai/solvi-large (preview)

A 396M-parameter ModernBERT-large cross-encoder that answers typed questions about a text or a JSON state for solvi: one option, several options, ordered scores, yes / no, not stated, spans, evidence quotes, rankings and numbers (bins), with a calibrated confidence and an act / escalate signal per answer. Same format as the 150M solvi-base (solvi_decide v2, subformat l14g typed v2); this is the larger, stronger variant.

Honest summary. It is clearly better than our previous deciders on every test, and better than GLiNER2.5-Decide on typed questions over JSON states — but it does not beat GLiNER2.5-Decide on zero-shot choice questions (Fast Decisions dev 59.4% vs 62.9%). With 64 labelled examples per head it reaches 63.0%. It is also slower: 137 ms per question on a CPU (ONNX fp16).

Results (held-out tests; never used for training or checkpoint selection)

Criteria were registered before training (research log L19; 4 of 10 met for this model — the misses are under Limitations). The Fast Decisions numbers are on the public dev split (100 rows × 17 domains) in our harness; Fastino's official scores (GLiNER2.5-Decide 60.1, …) are on a hidden test split and are not directly comparable.

Choice questions — Fast Decisions (dev, macro over domains):

model params zero-shot + bias correction, unlabelled (Zc) + 32 labelled (Sc) + 64 labelled (Sc)
GLiNER2.5-Decide (fastino) 340M 62.9% — — —
solvi-large (this model) 396M 59.4% 60.5% 62.0% 63.0%
solvi-base (distilled student) 150M 56.3% 58.2% 60.2% 60.4%
L14g (previous base, not published) 150M 56.2% 57.4% 58.8% —
solvi L14d (previous clean decider) 150M 57.3% 59.5% 61.4% 61.9%
knowledgator/gliformer-large-v1 (our run) 576M 43.9% — — —
Laya (our earlier run) — 48.9% — — —

Single-label heads alone: 63.4%; exact multi-label sets: 26.3%.

Typed questions over JSON states — typed-decisions test (2000 questions):

model zero-shot fitted per process, 32 / 300 examples
solvi-large 54.5% 63.1% / 64.9%
GLiNER2.5-Decide (our run) 50.3% —
L14g (previous base) 45.1% 58.8% / 61.3%
Laya base 36.0% fine-tuned on this dataset: 76.7%

typed-decisions gold comes from a ~4B teacher's judgement, so "accuracy" here means agreement with that teacher.

New kinds on real text:

test measure result
ContractNLI test (in-distribution: ContractNLI train was used for training) yes / no / not stated 88.8%
ContractNLI the evidence quote supports the answer (entailed / contradicted pairs) 62.0% (previous base L14g: 23%)
Taskmaster-2, held-out slots span exact / token F1 70.2% / 0.83
Taskmaster-2 "not stated" (noisy gold) 35.5%

Synthetic states with 9 held-out schemas: all kinds 97.9% (prose rendering 98.6%); yes / no / not stated 98.9%; span exact 99.4%; evidence quotes literally correct 100% and supporting 99.5%; ranking NDCG@3 0.983; number median bin 95.8%.

Act / escalate signal (AUROC of "the answer is right"): ContractNLI 0.85, Fast Decisions 0.76, Taskmaster-2 0.74, typed-decisions 0.67, synthetic 0.93. The shipped act calibrator is fitted on development data and is over-confident on new real text — set the threshold on your data with act_guard (below).

Calibration (ECE): ContractNLI 0.026, Taskmaster-2 0.062, synthetic ≤ 0.02 per kind (rank 0.21), typed-decisions 0.21, Fast Decisions 0.21 (fit on your task before trusting zero-shot confidences).

Speed (CPU, 4 threads, idle machine): ONNX fp16, one question with 10 options and ~20 words: 137 ms (fp32 128 ms); a question over a long JSON state: 450–740 ms. ONNX = PyTorch on 300 of 300 Fast Decisions heads.

Several questions per pass (block layout) — answers agree with one-question-per-pass only 90.8% (typed-decisions) / 90.2% (Fast Decisions) / 97.5% (synthetic) of the time, so solvi_decide.json ships with multi_question.enabled = false. The block ONNX export (onnx/model_block_fp16.onnx) is included for experiments (same answer as PyTorch 99.3%).

Training

  • Initial weights: answerdotai/ModernBERT-large (Apache-2.0).
  • Recipe: 200 minutes on one GPU per run (≈ 9.7k updates, effective batch 128, layer-wise LR decay 0.92), a first stage with 75% classification data, then the full mix; 35% of batches in the block layout with a block/full consistency loss. This model is the weight average of the best and the last checkpoint of that run, chosen by a development metric (8 permissive held-out classification sets + development splits of the training sources) — never by the tests above. Three sibling runs (a capacity-aware single-stage mix, a run without GLiNER distillation, a clean-data annealing of the second run) scored lower on that metric; averaging across runs made it worse.
  • Teachers (labels only on the license-clean texts below; never on test data), all Apache-2.0: Qwen/Qwen2.5-32B-Instruct, Qwen/Qwen2.5-14B-Instruct, mistralai/Mistral-Small-24B-Instruct-2501, fastino/GLiNER2.5-Decide. Soft targets = vote shares; questions without ≥ 2 agreeing votes were dropped. On heads with real gold the consensus was right 87.5% of the time (GLiNER alone 73.3%).
  • Data agent: Qwen2.5-32B-Instruct generated 22.7k texts with classification schemas over 76 business / content domains × 28 task types × 26 layouts; each head was verified by Qwen2.5-14B and Mistral-Small (kept at ≥ 2 of 3). The domain matrix was written from scratch and overlaps the public category descriptions of Fast Decisions; no Fast Decisions row was used, and all generated texts were de-duplicated against the Fast Decisions dev split (0 near-duplicates).
  • De-duplication: exact + MinHash (Jaccard ≥ 0.5) against Fast Decisions dev, typed-decisions test, ContractNLI test and Taskmaster-2 test texts; 216 ContractNLI-train windows too close to test contracts were removed.

Training data and attribution

The full manifest with licenses, URLs and counts is LICENSES.json. Policy: permissive, public-domain and share-alike data (CC BY-SA, with attribution here); excluded: non-commercial, research-only, unlicensed, or terms that forbid training.

source license attribution
solvi synthetic corpora (operational texts, typed states, judgment states; own generators, no external text) Apache-2.0 solvi authors
SQuAD 2.0 (train) CC BY-SA 4.0 Pranav Rajpurkar, Robin Jia, Percy Liang: "Know What You Don't Know: Unanswerable Questions for SQuAD", ACL 2018
BoolQ (train) CC BY-SA 3.0 Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, Kristina Toutanova: "BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions", NAACL 2019
CFPB Consumer Complaint Database CC0-1.0 Consumer Financial Protection Bureau
English Wikinews (2023-07-28) CC BY 2.5 Wikinews contributors, https://en.wikinews.org
arXiv abstracts (metadata) CC0-1.0 arXiv.org
Taskmaster-1 CC BY 4.0 Google Research
MultiWOZ 2.2 MIT Budzianowski et al.
Amazon polarity Apache-2.0 Zhang et al.; McAuley & Leskovec
GoEmotions Apache-2.0 Google Research
LEDGAR, UNFAIR-ToS (LexGLUE) CC BY 4.0 Tuggener et al.; Lippi et al.; Chalkidis et al.
SMS Spam Collection CC BY 4.0 Almeida & Hidalgo (UCI)
Toxic conversations 50k (Civil Comments / Jigsaw) CC BY 4.0 Jigsaw / Conversation AI; MTEB
deepset prompt-injections Apache-2.0 deepset
CLINC150 CC BY 3.0 Larson et al.
Banking77 CC BY 4.0 PolyAI
MASSIVE (en-US) CC BY 4.0 Amazon Science
SNIPS NLU Apache-2.0 Snips
Bitext customer-support, retail-banking, travel, insurance CDLA-Sharing-1.0 Bitext Innovations

| VitaminC (includes FEVER-based claims) | CC BY-SA 3.0 | Tal Schuster, Adam Fisch, Regina Barzilay: "Get Your Vitamin C! Robust Fact Verification with Contrastive Evidence", NAACL 2021; Wikipedia contributors | | ContractNLI (train split only) | CC BY 4.0 | Yuta Koreeda, Christopher D. Manning, "ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts", Findings of EMNLP 2021; Hitachi America, Ltd. | | LLM-generated texts and labels (Qwen2.5-32B / 14B-Instruct, Mistral-Small-24B-Instruct-2501) | Apache-2.0 (model licenses) | Qwen team (Alibaba Cloud); Mistral AI |

SQuAD 2.0, BoolQ and VitaminC are share-alike datasets: this card is their attribution. The model weights are released under Apache-2.0; if you redistribute the datasets themselves, their own licenses apply.

Checkpoint selection only (not trained on): DBpedia-14 (CC BY-SA 3.0), HuffPost News Category (CC BY 4.0), zeroshot twitter-financial-news topic / sentiment (MIT), Civil Comments (CC0), poem_sentiment (CC BY 4.0), LexGLUE SCOTUS (CC BY 4.0), Amazon counterfactual (CC BY 4.0).

Independent benchmarks (as released, zero-shot)

benchmark solvi-large solvi-base others (their published numbers or the benchmark's leaderboard)
jabr classifier-benchmark v2 (49 tasks, 866 cases), macro 0.694 0.602 Jev 0.966, GLiNER2.5-Decide 0.739, Von 0.720, GLiNER2 0.684, Laya 0.583
decision-models-under-pressure: accuracy with 128 options 45.1% 50.8% Jev 60.0%, best other open model 41.0%, Laya 38.5%
same: answers changed by reordering 64 options 40.9% as listed · 0.5% in solvi ≥ 0.5.1 (sorted order by default) 24.1% as listed Jev 14.6%, Laya 49.4%
same: accuracy lost to near-duplicate distractors (64 options) 0.260 0.247 Jev 0.105, Laya 0.345

Yes / no questions are the weak spot on jabr: the model ranks them reasonably but says "yes" 36% of the time where the gold labels have 50% — calibrate the threshold on your data (act_guard, fit).

Escalation with a guarantee (solvi ≥ 0.5.1)

The act threshold shipped in solvi_decide.json was fitted on development data and does not keep its promise on new real text. Calibrate on a few hundred labelled examples of your own stream instead:

part.act_guard(examples, risk=0.10)   # P(answered alone and wrong) ≤ 10% of all questions, for inputs like the examples

Measured with 300 calibration examples per data set and 200 random splits (the rest of the set is the test). "Answered" is the share decided without a person, "error" the error among those, "risk" the share of all questions answered alone and wrong — the number the guarantee is about:

data set shipped "10%" threshold: answered / error act_guard(risk=0.10): answered / error / risk act_guard(risk=0.05)
typed-decisions 78% / 39% 32% / 30% / 9.8% 18% / 27% / 5.0%
Taskmaster-2 81% / 32% 50% / 20% / 10.0% 34% / 14% / 4.9%
ContractNLI 98% / 10% 97% / 10% / 9.6% 85% / 6% / 5.1%
JSON, 9 held-out schemas 100% / 2% 99.6% / 1.8% / 1.8% 99.7% / 1.9% / 1.9%

The guarantee holds on every set; how much can be automated depends on how hard the questions are. It holds for inputs like the calibration examples, not under a shift of domain: recalibrate when your inputs change.

Evaluation only — never trained on: Fast Decisions (fastino, dev split), typed-decisions test, ContractNLI test, Taskmaster-2 held-out domains.

Limitations

  • Not better than GLiNER2.5-Decide on zero-shot choice questions (59.4% vs 62.9% on Fast Decisions dev). It reaches GLiNER's zero-shot level only after fitting on ~64 labelled examples per head.
  • Slower than 100 ms on a CPU (137 ms per short question, ONNX fp16); use solvi-base (50 ms) where latency matters.
  • The act / escalate calibrator is over-confident on new real text (at the "≤ 10% error" threshold fitted on development data the real error was 32–39% on Fast Decisions, typed-decisions and Taskmaster-2; 10% on ContractNLI). Use act_guard on your data (solvi ≥ 0.5.1): the risk it promises held on every set we measured.
  • Several questions per pass disagree with one-question-per-pass in ~9% of answers on real states — disabled by default.
  • ContractNLI numbers are in-distribution (train split used); evidence on other document types is weaker.
  • "Not stated" in dialogues is under-recognised (35.5% on Taskmaster-2); rank confidences are poorly calibrated (ECE 0.21).
  • English only; inputs longer than 512 tokens (1024 in the block layout) are truncated — only the text, never the question.
  • Not for medical, legal or credit decisions on its own; in solvi those rules belong in hard checks with escalation to a person.

Files

model.safetensors (bf16), onnx/model_fp16.onnx (full layout: input_ids, attention_mask → logits [B, L, 6]), onnx/model_block_fp16.onnx (block layout: input_ids, position_ids, full_attention_mask, sliding_attention_mask), onnx/parity.json, config.json, tokenizer.json, tokenizer_config.json, solvi_decide.json (capabilities, temperatures, act calibrator, multi_question.enabled = false), l14g_format.py and l14f_format.py (the exact input / output format), LICENSES.json (training data manifest).

Citation

@software{solvi,
  title  = {solvi: verifiable decision systems from catalogs of functions and checks},
  author = {mxkuzn and solvi contributors},
  year   = {2026},
  url    = {https://github.com/solvi-ai/solvi}
}
Downloads last month
26
Safetensors
Model size
0.4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for solvi-ai/solvi-large

Quantized
(23)
this model
Quantizations
1 model

Datasets used to train solvi-ai/solvi-large