solvi-ai/solvi-base (preview)

A 150M-parameter ModernBERT-base cross-encoder for solvi typed questions (same format and answer kinds as solvi-large), distilled from solvi-large for CPU / browser use: 50 ms per question on a CPU (ONNX fp16).

Honest summary. On typed questions over JSON states it matches the large model (typed-decisions 54.5%), and its contract evidence is much better than the previous base model (46% vs 23% supporting quotes). On zero-shot choice questions it is not better than the previous base models (Fast Decisions dev 56.3%; L14d 57.3%, GLiNER2.5-Decide 62.9%).

Results (held-out tests; never used for training or checkpoint selection)

Fast Decisions numbers are on the public dev split in our harness (Fastino's official numbers are on a hidden test split).

test solvi-base (this) solvi-large (teacher) L14g (previous base)
Fast Decisions dev, zero-shot / Zc / Sc k=32 56.3% / 58.2% / 60.2% 59.4% / 60.5% / 62.0% 56.2% / 57.4% / 58.8%
typed-decisions, zero-shot / k=300 54.5% / 64.0% 54.5% / 64.9% 45.1% / 61.3%
ContractNLI (in-distribution), three answers 87.3% 88.8% 56.0%
ContractNLI, evidence supports the answer 46.4% 62.0% 23.3%
Taskmaster-2 held-out slots, span F1 / exact 0.77 / 66.1% 0.83 / 70.2% 0.80 / 67.7%
synthetic held-out schemas, all kinds 94.9% 97.9% 94.4%
act AUROC: Fast Decisions / typed-decisions / ContractNLI / Taskmaster-2 0.74 / 0.60 / 0.86 / 0.64 0.76 / 0.67 / 0.85 / 0.74 0.71 / 0.57 / 0.63 / 0.64
ECE: typed-decisions / ContractNLI / Taskmaster-2 0.12 / 0.02 / 0.11 0.21 / 0.03 / 0.06 0.19 / 0.04 / 0.10

For reference: GLiNER2.5-Decide 62.9% on Fast Decisions dev and 50.3% on typed-decisions; Laya base 36.0% on typed-decisions.

Speed (CPU, 4 threads): ONNX fp16, one question with 10 options and ~20 words: 50 ms (fp32 46 ms); a question over a long JSON state ~170–270 ms. Block layout (several questions per pass) agrees with one question per pass only 90.2% on typed-decisions, so multi_question.enabled = false; the block ONNX export is included for experiments.

Training

  • Initial weights: L14g (our previous base, ModernBERT-base lineage, Apache-2.0).
  • Distillation: 65 minutes on one GPU (≈ 9.3k updates, batch 64) on the same mix as solvi-large; every target is 0.5 × the data label + 0.5 × solvi-large's probabilities. Checkpoint chosen on the same development metric as the large model, never on the tests.
  • Data and teachers: as solvi-large (see its card): votes of Qwen2.5-32B / 14B-Instruct, Mistral-Small-24B-Instruct-2501 and GLiNER2.5-Decide (all Apache-2.0), a Qwen2.5-32B data agent verified by two other models, ContractNLI train, VitaminC, SQuAD 2.0, BoolQ, synthetic typed states; de-duplicated against every test set.

Training data and attribution

The full manifest with licenses, URLs and counts is LICENSES.json. Policy: permissive, public-domain and share-alike data (CC BY-SA, with attribution here); excluded: non-commercial, research-only, unlicensed, or terms that forbid training.

source license attribution
solvi synthetic corpora (operational texts, typed states, judgment states; own generators, no external text) Apache-2.0 solvi authors
SQuAD 2.0 (train) CC BY-SA 4.0 Pranav Rajpurkar, Robin Jia, Percy Liang: "Know What You Don't Know: Unanswerable Questions for SQuAD", ACL 2018
BoolQ (train) CC BY-SA 3.0 Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, Kristina Toutanova: "BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions", NAACL 2019
CFPB Consumer Complaint Database CC0-1.0 Consumer Financial Protection Bureau
English Wikinews (2023-07-28) CC BY 2.5 Wikinews contributors, https://en.wikinews.org
arXiv abstracts (metadata) CC0-1.0 arXiv.org
Taskmaster-1 CC BY 4.0 Google Research
MultiWOZ 2.2 MIT Budzianowski et al.
Amazon polarity Apache-2.0 Zhang et al.; McAuley & Leskovec
GoEmotions Apache-2.0 Google Research
LEDGAR, UNFAIR-ToS (LexGLUE) CC BY 4.0 Tuggener et al.; Lippi et al.; Chalkidis et al.
SMS Spam Collection CC BY 4.0 Almeida & Hidalgo (UCI)
Toxic conversations 50k (Civil Comments / Jigsaw) CC BY 4.0 Jigsaw / Conversation AI; MTEB
deepset prompt-injections Apache-2.0 deepset
CLINC150 CC BY 3.0 Larson et al.
Banking77 CC BY 4.0 PolyAI
MASSIVE (en-US) CC BY 4.0 Amazon Science
SNIPS NLU Apache-2.0 Snips
Bitext customer-support, retail-banking, travel, insurance CDLA-Sharing-1.0 Bitext Innovations

| VitaminC (includes FEVER-based claims) | CC BY-SA 3.0 | Tal Schuster, Adam Fisch, Regina Barzilay: "Get Your Vitamin C! Robust Fact Verification with Contrastive Evidence", NAACL 2021; Wikipedia contributors | | ContractNLI (train split only) | CC BY 4.0 | Yuta Koreeda, Christopher D. Manning, "ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts", Findings of EMNLP 2021; Hitachi America, Ltd. | | LLM-generated texts and labels (Qwen2.5-32B / 14B-Instruct, Mistral-Small-24B-Instruct-2501) | Apache-2.0 (model licenses) | Qwen team (Alibaba Cloud); Mistral AI |

SQuAD 2.0, BoolQ and VitaminC are share-alike datasets: this card is their attribution. The model weights are released under Apache-2.0; if you redistribute the datasets themselves, their own licenses apply.

Checkpoint selection only (not trained on): DBpedia-14 (CC BY-SA 3.0), HuffPost News Category (CC BY 4.0), zeroshot twitter-financial-news topic / sentiment (MIT), Civil Comments (CC0), poem_sentiment (CC BY 4.0), LexGLUE SCOTUS (CC BY 4.0), Amazon counterfactual (CC BY 4.0).

Independent benchmarks (as released, zero-shot)

benchmark solvi-large solvi-base others (their published numbers or the benchmark's leaderboard)
jabr classifier-benchmark v2 (49 tasks, 866 cases), macro 0.694 0.602 Jev 0.966, GLiNER2.5-Decide 0.739, Von 0.720, GLiNER2 0.684, Laya 0.583
decision-models-under-pressure: accuracy with 128 options 45.1% 50.8% Jev 60.0%, best other open model 41.0%, Laya 38.5%
same: answers changed by reordering 64 options 40.9% as listed · 0.5% in solvi ≥ 0.5.1 (sorted order by default) 24.1% as listed Jev 14.6%, Laya 49.4%
same: accuracy lost to near-duplicate distractors (64 options) 0.260 0.247 Jev 0.105, Laya 0.345

Yes / no questions are the weak spot on jabr: the model ranks them reasonably but says "yes" 36% of the time where the gold labels have 50% — calibrate the threshold on your data (act_guard, fit).

Escalation with a guarantee (solvi ≥ 0.5.1)

The act threshold shipped in solvi_decide.json was fitted on development data and does not keep its promise on new real text. Calibrate on a few hundred labelled examples of your own stream instead:

part.act_guard(examples, risk=0.10)   # P(answered alone and wrong) ≤ 10% of all questions, for inputs like the examples

Measured with 300 calibration examples per data set and 200 random splits (the rest of the set is the test). "Answered" is the share decided without a person, "error" the error among those, "risk" the share of all questions answered alone and wrong — the number the guarantee is about:

data set shipped "10%" threshold: answered / error act_guard(risk=0.10): answered / error / risk act_guard(risk=0.05)
typed-decisions 66% / 40% 27% / 36% / 9.9% 16% / 30% / 5.0%
Taskmaster-2 84% / 37% 39% / 25% / 10.1% 25% / 20% / 5.0%
ContractNLI 95% / 11% 93% / 10% / 9.6% 80% / 6% / 4.8%
JSON, 9 held-out schemas 97% / 3% 99.7% / 4.8% / 4.8% 99.0% / 4.2% / 4.2%

The guarantee holds on every set; how much can be automated depends on how hard the questions are. It holds for inputs like the calibration examples, not under a shift of domain: recalibrate when your inputs change.

Evaluation only — never trained on: Fast Decisions (fastino, dev split), typed-decisions test, ContractNLI test, Taskmaster-2 held-out domains.

Limitations

  • Zero-shot choice questions are not improved over the previous base models (Fast Decisions dev 56.3%; GLiNER2.5-Decide 62.9%). Fit on 30–60 examples of your task (Sc k=32: 60.2%).
  • Contract evidence supports the answer less often than the large model (46% vs 62%); act AUROC on typed-decisions 0.60.
  • The act / escalate thresholds shipped with the model are indicative only: fitted on development data, they are over-confident on new real text (see solvi-large). For a guaranteed risk, calibrate on your own labelled stream with act_guard (solvi ≥ 0.5.1, table above).
  • Several questions per pass disagree with one-question-per-pass in ~10% of answers on real states — disabled by default.
  • English only; truncation beyond 512 tokens (1024 in block layout). Not for medical, legal or credit decisions on its own.

Files

model.safetensors (bf16), onnx/model_fp16.onnx (full layout: input_ids, attention_mask → logits [B, L, 6]), onnx/model_block_fp16.onnx (block layout: input_ids, position_ids, full_attention_mask, sliding_attention_mask), onnx/parity.json, config.json, tokenizer.json, tokenizer_config.json, solvi_decide.json (capabilities, temperatures, act calibrator, multi_question.enabled = false), l14g_format.py and l14f_format.py (the exact input / output format), LICENSES.json (training data manifest).

Citation

@software{solvi,
  title  = {solvi: verifiable decision systems from catalogs of functions and checks},
  author = {mxkuzn and solvi contributors},
  year   = {2026},
  url    = {https://github.com/solvi-ai/solvi}
}
Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for solvi-ai/solvi-base

Quantized
(76)
this model

Datasets used to train solvi-ai/solvi-base