solvi-ai/solvi-large-long (preview)
solvi-large, fine-tuned to read documents up to 8,192 tokens whole. It
backs solvi's opt-in long="full" mode (solvi ≥ 0.7.0). The architecture (ModernBERT-large, 396M), the answer kinds and the
format (solvi_decide v2, subformat l14g typed v2) are the same as solvi-large. The one addition is "max_len_long": 8192
in solvi_decide.json; max_len stays 512, so short inputs are read exactly as before.
Honest summary.
- Long documents. On 4–8k-token contracts and reports, reading whole is right 84.6% of the time. solvi-large with
long="retrieve"gets 73.2%, and truncating at 512 tokens gets 43%. The evidence quote supports the answer in 76% of cases, against 50% for solvi-large with retrieve. - Escalation on long contracts. At the same 10% risk it answers on its own 95% of the time, against 70% for solvi-large with retrieve.
- Short inputs. The standard benchmarks stay within one point of solvi-large, and evidence quotes on contract windows are much better (87% vs 62%).
- What it costs, and what it missed.
- It is a GPU model. On a CPU, reading a whole 4k-token text costs about 12× a 512-token pass, an 8k text about 31×.
- It failed one of its own release checks: the act (escalation) signal on short contract windows is weaker than solvi-large's (AUROC 0.822 vs 0.854, −0.032 against an allowed −0.02). Details and the decision we took are below.
- Keep solvi-large as your default model; use this one where documents are long.
When to use it
| your inputs | use |
|---|---|
| short texts and JSON states (≤ 512 tokens) | solvi-large |
| long contracts and reports, on a GPU | this model with long="full" |
| long documents, on a CPU | this model with long="retrieve" and max_len=2048 (about 3× a 512-token pass), see Speed |
| long documents, CPU, cheapest | solvi-large with long="retrieve" (≈ 1× a 512-token pass) |
from solvi.decide import DecideModel
m = DecideModel.load("solvi-ai/solvi-large-long", device="cuda") # needs solvi >= 0.7.0
m.long_len # 8192
part = m.decision("assignment", "Either party may assign the agreement without consent.", "contract",
type=bool, unknown=True, evidence=True, long="full")
d = part(contract=contract_text)
d.value # True / False / solvi.Unknown ("not stated")
d.evidence # quotes: contract_text[q.start:q.end] == q.value
d.extra["long"] # {"mode": "full", "tokens": 5234, "max_len": 8192}; past 8192 tokens: + "fallback": "retrieve"
How long="full" reads a text.
- A text that fits 512 tokens is decided as before, in one ordinary pass.
- A longer one is read whole, in one pass of up to 8,192 tokens.
- Beyond 8,192 tokens, solvi falls back to retrieve within 8,192 tokens: sections of about 170 tokens, chosen by BM25
(
top_k=None, the default, picks how many).
Recommended settings.
- Keep
top_k=None. An explicittop_konly matters for texts longer than 8k tokens, or withlong="retrieve". - Calibrate the act threshold on your own long documents (
calibrate_for/act_guard). The shipped act calibrator is fitted on short development data. multi_questionstays disabled, as in solvi-large: one question per pass.
Results
All tests below are held out: none was used for training or checkpoint selection. Release criteria were registered before training.
Long documents (4k–8k tokens)
Tests.
- ContractNLI test: 121 NDAs × 17 hypotheses, yes / no / not stated, gold evidence spans from the annotators.
- This is in-distribution by genre: ContractNLI train contracts were used for training.
- 78 held-out long documents: CUAD contracts, BillSum bills and GovReport reports never used in training.
- Questions: yes / no / not stated, choice, number and span.
- Labels come from the consensus of 2 of 3 open teacher models, which agreed with real gold 82–86% of the time.
Each question is asked on a window of the document that contains all of its evidence (1094 questions at 4k and 8k tokens, 86 documents).
| model and mode | accuracy, 4k–8k | evidence quote supports the answer |
|---|---|---|
solvi-large, truncate at 512 (long=None) |
43.2% | 11.1% |
solvi-large, long="retrieve" (3 sections, 512 tokens) |
73.2% | 49.5% |
| solvi-large, read whole (not trained for it) | 73.7% | 34.7% |
this model, long="retrieve" (3 sections, 512 tokens) |
76.9% | 68.0% |
this model, long="retrieve", max_len=2048 (12 sections) |
85.3% | 73.0% |
this model, long="full" |
84.6% | 76.3% |
By length, long="full":
| length | 0.5k | 1k | 2k | 4k | 8k |
|---|---|---|---|---|---|
| accuracy | 89.7% | 87.7% | 86.8% | 85.1% | 80.5% (118 questions) |
By test and question type (4k–8k):
- By test: ContractNLI 86.5%, the held-out documents 82.9%.
- Yes / no / not stated gains the most: 85.6% whole, against 74.1% with this model's retrieve and 71.3% with solvi-large's. To say "the contract does not say this", the model has to see the whole document.
- Choice and number questions: retrieve is as good or better (number: 76.9% with retrieve vs 73.8% whole), because the answer sits in one place and BM25 finds it.
Escalation on long documents (2k–8k tokens)
How it was measured.
- Rows: 3,119 questions at 2k, 4k and 8k tokens (ContractNLI: 1,890 on 81 contracts; held-out documents: 1,229 on 78).
- Calibration: the act calibrator (logistic regression on the model's confidence features) is fitted on half of the documents and tested on the other half, over 200 random splits.
- Threshold: a risk-controlled act threshold, P(auto-answer and wrong) ≤ 0.10.
| test | model | accuracy | AUROC act | answered on its own at risk 0.10 | realised risk | ECE |
|---|---|---|---|---|---|---|
| ContractNLI | solvi-large + retrieve | 75.6% | 0.784 | 70.0% | 0.098 | 0.136 |
| ContractNLI | this model, full | 87.9% | 0.826 | 95.5% | 0.101 | 0.040 |
| held-out documents | solvi-large + retrieve | 73.3% | 0.799 | 71.6% | 0.108 | 0.109 |
| held-out documents | this model, full | 83.4% | 0.788 | 85.6% | 0.100 | 0.059 |
On the held-out documents its act ranking is slightly weaker than solvi-large with retrieve (0.788 vs 0.799; on the 4–8k ones alone 0.756 vs 0.799). It still answers on its own more often at the same risk, because it is right more often.
Short inputs: the solvi-large benchmarks
| test | solvi-large | this model |
|---|---|---|
| Fast Decisions dev, zero-shot (macro over domains) | 59.4% | 59.1% |
| typed-decisions test, zero-shot | 54.5% | 54.4% |
| synthetic states, 9 held-out schemas (all kinds) | 97.9% | 98.0% |
| ContractNLI windows, yes / no / not stated (in-distribution) | 88.8% | 91.4% |
| ContractNLI windows, evidence quote supports the answer | 62.0% | 87.3% |
| Taskmaster-2 held-out slots, span token F1 | 0.829 | 0.838 |
| jabr/classifier-benchmark v2, macro | 0.694 | 0.686 |
| decision-models-under-pressure, canonical order, 16 / 64 options | 82.2% / 63.1% | 81.9% / 62.2% |
| act AUROC, short inputs: Fast Decisions / typed-decisions / Taskmaster-2 | 0.762 / 0.668 / 0.743 | 0.758 / 0.663 / 0.751 |
| act AUROC, short ContractNLI windows | 0.854 | 0.822 |
The release check it failed
Before training we registered release criteria for a long-input model. One of them compared the act signal with solvi-large's on short test windows: each set's AUROC could drop by at most 0.02. On short ContractNLI windows (≤ 280 words) this model drops by 0.032 (0.854 → 0.822), so it fails that check. The other short-input checks pass (average act AUROC −0.008 against an allowed −0.01; every benchmark within −1.0 point).
The decision. This model is only offered for the opt-in long="full" mode, where texts longer than 512 tokens are read
whole. On 2026-09-29, before measuring, we decided that for this mode the short-window act check is replaced by the same
check on long documents, where the mode actually runs (the escalation table above; the pass rule: act AUROC at least
solvi-large + retrieve − 0.02 on each test, at least as many auto-answers at risk 0.10, realised risk ≤ 0.11). The model
passes it. The default model's release rule is unchanged. The failed number is stated here so you can judge it: if your
contract inputs are short, use solvi-large.
Speed and the CPU alternative
| hardware | measurement | time |
|---|---|---|
| CPU, 4 threads, fp32 (relative to a 512-token pass) | whole 2k / 4k / 8k-token text | about 5× / 12× / 31× |
| CPU, same | on a typical laptop CPU | ≈ 1.6 s per question at 4k, ≈ 4 s at 8k |
| CPU, same | long="retrieve", max_len=2048 (reads at most 2,048 tokens) |
about 3× |
| GPU, RTX 3060 Laptop, 6 GB, bf16 (same architecture) | 12 questions on 8k-token texts | 15.7 s, 2.8 GB peak memory |
| GPU, RTX PRO 6000 | our whole long-document evaluation: 8,274 questions × 4 reading modes | 7 min |
On a CPU, load the model with a larger retrieve budget instead of reading whole:
m = DecideModel.load("solvi-ai/solvi-large-long", max_len=2048) # top_k=None: 12 sections of ≈ 170 tokens
part = m.decision("assignment", "...", "contract", type=bool, unknown=True, long="retrieve")
Measured on the same 1,094 questions at 4k–8k tokens, this reads at most 2,048 tokens per question and is as accurate as
reading whole: 85.3% vs 84.6% (difference +0.6 points, 95% interval −1.1 to +2.4, bootstrap over documents). Its quotes
support the answer a little less often (73.0% vs 76.3%). This budget helps only a model trained on long inputs: solvi-large
with max_len=2048 stays at 73%.
Training
- Initial weights: solvi-ai/solvi-large (Apache-2.0).
- Recipe: 50 minutes on one GPU (RTX PRO 6000, 96 GB), about 1,300 updates of 128 examples.
- 15% whole-document examples: a window of 1.5k–8k tokens (log-uniform) containing the evidence, read in one sequence
of up to 8,192 tokens.
- ContractNLI train contracts: 262 contracts, 4,454 hypotheses.
- CUAD / BillSum / GovReport train documents with open-teacher consensus: 179 documents, 1,485 questions.
- 85% replay of solvi-large's training mix.
- 15% whole-document examples: a window of 1.5k–8k tokens (log-uniform) containing the evidence, read in one sequence
of up to 8,192 tokens.
- Two changes over a plain fine-tune:
- Evidence spans longer than 40 tokens are clipped to their first 40 tokens in the pointer loss instead of being skipped; the pointer quotes at most 40 tokens. Nearly half of ContractNLI's evidence spans are that long. The same fine-tune without this change quoted supporting evidence on contract windows 53% of the time; with it, 87% (solvi-large: 62%).
- ContractNLI windows are replayed only for the act head, not in the main loss. This protects the short-input benchmarks.
- Checkpoint selection: long-document development accuracy (ContractNLI dev, read whole), among checkpoints whose solvi-large development metric did not drop by more than 0.01. Never the tests.
- Calibration: temperatures and the act calibrator are refitted for this model on development data, not on the tests.
Training data and attribution
Everything in solvi-large's card applies: its data, including the share-alike SQuAD 2.0, BoolQ and VitaminC; this card is their attribution too. Added for long documents:
| source | license | attribution |
|---|---|---|
| ContractNLI (train split, whole contracts) | CC BY 4.0 | Yuta Koreeda, Christopher D. Manning, "ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts", Findings of EMNLP 2021; Hitachi America, Ltd. |
| CUAD v1 (contracts outside the official test set) | CC BY 4.0 | The Atticus Project; Hendrycks et al., "CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review", 2021 |
| BillSum (train) | CC0-1.0 | FiscalNote (Kornilova & Eidelman, 2019); US government works |
| GovReport (train; the first 2k–7.5k tokens of a report) | CC BY 4.0 | Huang et al., "Efficient Attentions for Long Document Summarization", NAACL 2021; US GAO / CRS |
| Questions, labels and quotes for long documents: Qwen/Qwen3-235B-A22B-Instruct-2507, deepseek-ai/DeepSeek-V3.2, openai/gpt-oss-120b | Apache-2.0 / MIT / Apache-2.0 (model licenses allow training on outputs) | Qwen team (Alibaba Cloud); DeepSeek-AI; OpenAI |
The full manifest is LICENSES.json.
Evaluation only, never trained on:
- ContractNLI test;
- the 78 held-out long documents;
- Fast Decisions (dev);
- typed-decisions test;
- Taskmaster-2 held-out domains;
- jabr/classifier-benchmark v2;
- decision-models-under-pressure.
Limitations
- A GPU model. Reading 8k tokens whole on a CPU takes seconds per question. solvi warns once per model when
long="full"reads more than 2k tokens on a CPU. - The act signal on short contract windows is weaker than solvi-large's (AUROC 0.822 vs 0.854); this failed a release check (above). Calibrate on your data.
- Evaluation limits.
- ContractNLI results are in-distribution by genre.
- The 8k bucket is small: 118 questions.
- The held-out long documents are labelled by teacher-model consensus, not by human annotators.
- Evidence quotes are at most 40 tokens. A long clause is quoted by its beginning.
- Short-input benchmarks are within a point of solvi-large but slightly lower on Fast Decisions (−0.3), jabr (−0.8) and decision-models-under-pressure (−0.3 / −0.9).
- Past 8,192 tokens the model does not see the whole text: solvi retrieves sections within 8,192 tokens, and a question whose evidence BM25 misses can be answered "not stated".
- Inherited from solvi-large:
- English only.
- Not better than GLiNER2.5-Decide on zero-shot choice questions.
- Several questions per pass is disabled.
- Not for medical, legal or credit decisions on its own.
Files
model.safetensors(bf16).onnx/model_fp16.onnx: full layout, dynamic length: input_ids, attention_mask → logits [B, L, 6].onnx/model_block_fp16.onnx: block layout, for experiments.onnx/parity.json: ONNX vs PyTorch on short inputs and on 2k / 4k-token documents.config.json,tokenizer.json,tokenizer_config.json.solvi_decide.json: capabilities,max_len: 512,max_len_long: 8192, temperatures, act calibrator,multi_question.enabled = false.l14g_format.py,l14f_format.py: the reference input format.LICENSES.json.
Citation
@software{solvi,
title = {solvi: verifiable decision systems from catalogs of functions and checks},
author = {mxkuzn and solvi contributors},
year = {2026},
url = {https://github.com/solvi-ai/solvi}
}
- Downloads last month
- 11