solvi-ai/solvi-base (preview)
A 150M-parameter ModernBERT-base cross-encoder for solvi typed questions (same format and answer kinds as solvi-large), distilled from solvi-large for CPU / browser use: 50 ms per question on a CPU (ONNX fp16).
Honest summary. On typed questions over JSON states it matches the large model (typed-decisions 54.5%), and its contract evidence is much better than the previous base model (46% vs 23% supporting quotes). On zero-shot choice questions it is not better than the previous base models (Fast Decisions dev 56.3%; L14d 57.3%, GLiNER2.5-Decide 62.9%).
Results (held-out tests; never used for training or checkpoint selection)
Fast Decisions numbers are on the public dev split in our harness (Fastino's official numbers are on a hidden test split).
| test | solvi-base (this) | solvi-large (teacher) | L14g (previous base) |
|---|---|---|---|
| Fast Decisions dev, zero-shot / Zc / Sc k=32 | 56.3% / 58.2% / 60.2% | 59.4% / 60.5% / 62.0% | 56.2% / 57.4% / 58.8% |
| typed-decisions, zero-shot / k=300 | 54.5% / 64.0% | 54.5% / 64.9% | 45.1% / 61.3% |
| ContractNLI (in-distribution), three answers | 87.3% | 88.8% | 56.0% |
| ContractNLI, evidence supports the answer | 46.4% | 62.0% | 23.3% |
| Taskmaster-2 held-out slots, span F1 / exact | 0.77 / 66.1% | 0.83 / 70.2% | 0.80 / 67.7% |
| synthetic held-out schemas, all kinds | 94.9% | 97.9% | 94.4% |
| act AUROC: Fast Decisions / typed-decisions / ContractNLI / Taskmaster-2 | 0.74 / 0.60 / 0.86 / 0.64 | 0.76 / 0.67 / 0.85 / 0.74 | 0.71 / 0.57 / 0.63 / 0.64 |
| ECE: typed-decisions / ContractNLI / Taskmaster-2 | 0.12 / 0.02 / 0.11 | 0.21 / 0.03 / 0.06 | 0.19 / 0.04 / 0.10 |
For reference: GLiNER2.5-Decide 62.9% on Fast Decisions dev and 50.3% on typed-decisions; Laya base 36.0% on typed-decisions.
Speed (CPU, 4 threads): ONNX fp16, one question with 10 options and ~20 words: 50 ms (fp32 46 ms); a question over a
long JSON state ~170–270 ms. Block layout (several questions per pass) agrees with one question per pass only 90.2% on
typed-decisions, so multi_question.enabled = false; the block ONNX export is included for experiments.
Training
- Initial weights: L14g (our previous base, ModernBERT-base lineage, Apache-2.0).
- Distillation: 65 minutes on one GPU (≈ 9.3k updates, batch 64) on the same mix as solvi-large; every target is 0.5 × the data label + 0.5 × solvi-large's probabilities. Checkpoint chosen on the same development metric as the large model, never on the tests.
- Data and teachers: as solvi-large (see its card): votes of Qwen2.5-32B / 14B-Instruct, Mistral-Small-24B-Instruct-2501 and GLiNER2.5-Decide (all Apache-2.0), a Qwen2.5-32B data agent verified by two other models, ContractNLI train, VitaminC, SQuAD 2.0, BoolQ, synthetic typed states; de-duplicated against every test set.
Training data and attribution
The full manifest with licenses, URLs and counts is LICENSES.json. Policy: permissive, public-domain and share-alike data (CC BY-SA,
with attribution here); excluded: non-commercial, research-only, unlicensed, or terms that forbid training.
| source | license | attribution |
|---|---|---|
| solvi synthetic corpora (operational texts, typed states, judgment states; own generators, no external text) | Apache-2.0 | solvi authors |
| SQuAD 2.0 (train) | CC BY-SA 4.0 | Pranav Rajpurkar, Robin Jia, Percy Liang: "Know What You Don't Know: Unanswerable Questions for SQuAD", ACL 2018 |
| BoolQ (train) | CC BY-SA 3.0 | Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, Kristina Toutanova: "BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions", NAACL 2019 |
| CFPB Consumer Complaint Database | CC0-1.0 | Consumer Financial Protection Bureau |
| English Wikinews (2023-07-28) | CC BY 2.5 | Wikinews contributors, https://en.wikinews.org |
| arXiv abstracts (metadata) | CC0-1.0 | arXiv.org |
| Taskmaster-1 | CC BY 4.0 | Google Research |
| MultiWOZ 2.2 | MIT | Budzianowski et al. |
| Amazon polarity | Apache-2.0 | Zhang et al.; McAuley & Leskovec |
| GoEmotions | Apache-2.0 | Google Research |
| LEDGAR, UNFAIR-ToS (LexGLUE) | CC BY 4.0 | Tuggener et al.; Lippi et al.; Chalkidis et al. |
| SMS Spam Collection | CC BY 4.0 | Almeida & Hidalgo (UCI) |
| Toxic conversations 50k (Civil Comments / Jigsaw) | CC BY 4.0 | Jigsaw / Conversation AI; MTEB |
| deepset prompt-injections | Apache-2.0 | deepset |
| CLINC150 | CC BY 3.0 | Larson et al. |
| Banking77 | CC BY 4.0 | PolyAI |
| MASSIVE (en-US) | CC BY 4.0 | Amazon Science |
| SNIPS NLU | Apache-2.0 | Snips |
| Bitext customer-support, retail-banking, travel, insurance | CDLA-Sharing-1.0 | Bitext Innovations |
| VitaminC (includes FEVER-based claims) | CC BY-SA 3.0 | Tal Schuster, Adam Fisch, Regina Barzilay: "Get Your Vitamin C! Robust Fact Verification with Contrastive Evidence", NAACL 2021; Wikipedia contributors | | ContractNLI (train split only) | CC BY 4.0 | Yuta Koreeda, Christopher D. Manning, "ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts", Findings of EMNLP 2021; Hitachi America, Ltd. | | LLM-generated texts and labels (Qwen2.5-32B / 14B-Instruct, Mistral-Small-24B-Instruct-2501) | Apache-2.0 (model licenses) | Qwen team (Alibaba Cloud); Mistral AI |
SQuAD 2.0, BoolQ and VitaminC are share-alike datasets: this card is their attribution. The model weights are released under Apache-2.0; if you redistribute the datasets themselves, their own licenses apply.
Checkpoint selection only (not trained on): DBpedia-14 (CC BY-SA 3.0), HuffPost News Category (CC BY 4.0), zeroshot twitter-financial-news topic / sentiment (MIT), Civil Comments (CC0), poem_sentiment (CC BY 4.0), LexGLUE SCOTUS (CC BY 4.0), Amazon counterfactual (CC BY 4.0).
Independent benchmarks (as released, zero-shot)
| benchmark | solvi-large | solvi-base | others (their published numbers or the benchmark's leaderboard) |
|---|---|---|---|
| jabr classifier-benchmark v2 (49 tasks, 866 cases), macro | 0.694 | 0.602 | Jev 0.966, GLiNER2.5-Decide 0.739, Von 0.720, GLiNER2 0.684, Laya 0.583 |
| decision-models-under-pressure: accuracy with 128 options | 45.1% | 50.8% | Jev 60.0%, best other open model 41.0%, Laya 38.5% |
| same: answers changed by reordering 64 options | 40.9% as listed · 0.5% in solvi ≥ 0.5.1 (sorted order by default) | 24.1% as listed | Jev 14.6%, Laya 49.4% |
| same: accuracy lost to near-duplicate distractors (64 options) | 0.260 | 0.247 | Jev 0.105, Laya 0.345 |
Yes / no questions are the weak spot on jabr: the model ranks them reasonably but says "yes" 36% of the time where the gold
labels have 50% — calibrate the threshold on your data (act_guard, fit).
Escalation with a guarantee (solvi ≥ 0.5.1)
The act threshold shipped in solvi_decide.json was fitted on development data and does not keep its promise on new real
text. Calibrate on a few hundred labelled examples of your own stream instead:
part.act_guard(examples, risk=0.10) # P(answered alone and wrong) ≤ 10% of all questions, for inputs like the examples
Measured with 300 calibration examples per data set and 200 random splits (the rest of the set is the test). "Answered" is the share decided without a person, "error" the error among those, "risk" the share of all questions answered alone and wrong — the number the guarantee is about:
| data set | shipped "10%" threshold: answered / error | act_guard(risk=0.10): answered / error / risk |
act_guard(risk=0.05) |
|---|---|---|---|
| typed-decisions | 66% / 40% | 27% / 36% / 9.9% | 16% / 30% / 5.0% |
| Taskmaster-2 | 84% / 37% | 39% / 25% / 10.1% | 25% / 20% / 5.0% |
| ContractNLI | 95% / 11% | 93% / 10% / 9.6% | 80% / 6% / 4.8% |
| JSON, 9 held-out schemas | 97% / 3% | 99.7% / 4.8% / 4.8% | 99.0% / 4.2% / 4.2% |
The guarantee holds on every set; how much can be automated depends on how hard the questions are. It holds for inputs like the calibration examples, not under a shift of domain: recalibrate when your inputs change.
Evaluation only — never trained on: Fast Decisions (fastino, dev split), typed-decisions test, ContractNLI test, Taskmaster-2 held-out domains.
Limitations
- Zero-shot choice questions are not improved over the previous base models (Fast Decisions dev 56.3%; GLiNER2.5-Decide 62.9%). Fit on 30–60 examples of your task (Sc k=32: 60.2%).
- Contract evidence supports the answer less often than the large model (46% vs 62%); act AUROC on typed-decisions 0.60.
- The act / escalate thresholds shipped with the model are indicative only: fitted on development data, they are
over-confident on new real text (see solvi-large). For a guaranteed risk, calibrate on your own labelled stream
with
act_guard(solvi ≥ 0.5.1, table above). - Several questions per pass disagree with one-question-per-pass in ~10% of answers on real states — disabled by default.
- English only; truncation beyond 512 tokens (1024 in block layout). Not for medical, legal or credit decisions on its own.
Files
model.safetensors (bf16), onnx/model_fp16.onnx (full layout: input_ids, attention_mask → logits [B, L, 6]),
onnx/model_block_fp16.onnx (block layout: input_ids, position_ids, full_attention_mask, sliding_attention_mask),
onnx/parity.json, config.json, tokenizer.json, tokenizer_config.json, solvi_decide.json (capabilities, temperatures,
act calibrator, multi_question.enabled = false), l14g_format.py and l14f_format.py (the exact input / output format),
LICENSES.json (training data manifest).
Citation
@software{solvi,
title = {solvi: verifiable decision systems from catalogs of functions and checks},
author = {mxkuzn and solvi contributors},
year = {2026},
url = {https://github.com/solvi-ai/solvi}
}
- Downloads last month
- -
Model tree for solvi-ai/solvi-base
Base model
answerdotai/ModernBERT-base