solvi-ai/solvi-large-long (preview)

solvi-large, fine-tuned to read documents up to 8,192 tokens whole. It backs solvi's opt-in long="full" mode (solvi ≥ 0.7.0). The architecture (ModernBERT-large, 396M), the answer kinds and the format (solvi_decide v2, subformat l14g typed v2) are the same as solvi-large. The one addition is "max_len_long": 8192 in solvi_decide.json; max_len stays 512, so short inputs are read exactly as before.

Honest summary.

  • Long documents. On 4–8k-token contracts and reports, reading whole is right 84.6% of the time. solvi-large with long="retrieve" gets 73.2%, and truncating at 512 tokens gets 43%. The evidence quote supports the answer in 76% of cases, against 50% for solvi-large with retrieve.
  • Escalation on long contracts. At the same 10% risk it answers on its own 95% of the time, against 70% for solvi-large with retrieve.
  • Short inputs. The standard benchmarks stay within one point of solvi-large, and evidence quotes on contract windows are much better (87% vs 62%).
  • What it costs, and what it missed.
    • It is a GPU model. On a CPU, reading a whole 4k-token text costs about 12× a 512-token pass, an 8k text about 31×.
    • It failed one of its own release checks: the act (escalation) signal on short contract windows is weaker than solvi-large's (AUROC 0.822 vs 0.854, −0.032 against an allowed −0.02). Details and the decision we took are below.
    • Keep solvi-large as your default model; use this one where documents are long.

When to use it

your inputs use
short texts and JSON states (≤ 512 tokens) solvi-large
long contracts and reports, on a GPU this model with long="full"
long documents, on a CPU this model with long="retrieve" and max_len=2048 (about 3× a 512-token pass), see Speed
long documents, CPU, cheapest solvi-large with long="retrieve" (≈ 1× a 512-token pass)
from solvi.decide import DecideModel

m = DecideModel.load("solvi-ai/solvi-large-long", device="cuda")      # needs solvi >= 0.7.0
m.long_len                                                              # 8192
part = m.decision("assignment", "Either party may assign the agreement without consent.", "contract",
                  type=bool, unknown=True, evidence=True, long="full")
d = part(contract=contract_text)
d.value            # True / False / solvi.Unknown ("not stated")
d.evidence         # quotes: contract_text[q.start:q.end] == q.value
d.extra["long"]    # {"mode": "full", "tokens": 5234, "max_len": 8192}; past 8192 tokens: + "fallback": "retrieve"

How long="full" reads a text.

  • A text that fits 512 tokens is decided as before, in one ordinary pass.
  • A longer one is read whole, in one pass of up to 8,192 tokens.
  • Beyond 8,192 tokens, solvi falls back to retrieve within 8,192 tokens: sections of about 170 tokens, chosen by BM25 (top_k=None, the default, picks how many).

Recommended settings.

  • Keep top_k=None. An explicit top_k only matters for texts longer than 8k tokens, or with long="retrieve".
  • Calibrate the act threshold on your own long documents (calibrate_for / act_guard). The shipped act calibrator is fitted on short development data.
  • multi_question stays disabled, as in solvi-large: one question per pass.

Results

All tests below are held out: none was used for training or checkpoint selection. Release criteria were registered before training.

Long documents (4k–8k tokens)

Tests.

  • ContractNLI test: 121 NDAs × 17 hypotheses, yes / no / not stated, gold evidence spans from the annotators.
    • This is in-distribution by genre: ContractNLI train contracts were used for training.
  • 78 held-out long documents: CUAD contracts, BillSum bills and GovReport reports never used in training.
    • Questions: yes / no / not stated, choice, number and span.
    • Labels come from the consensus of 2 of 3 open teacher models, which agreed with real gold 82–86% of the time.

Each question is asked on a window of the document that contains all of its evidence (1094 questions at 4k and 8k tokens, 86 documents).

model and mode accuracy, 4k–8k evidence quote supports the answer
solvi-large, truncate at 512 (long=None) 43.2% 11.1%
solvi-large, long="retrieve" (3 sections, 512 tokens) 73.2% 49.5%
solvi-large, read whole (not trained for it) 73.7% 34.7%
this model, long="retrieve" (3 sections, 512 tokens) 76.9% 68.0%
this model, long="retrieve", max_len=2048 (12 sections) 85.3% 73.0%
this model, long="full" 84.6% 76.3%

By length, long="full":

length 0.5k 1k 2k 4k 8k
accuracy 89.7% 87.7% 86.8% 85.1% 80.5% (118 questions)

By test and question type (4k–8k):

  • By test: ContractNLI 86.5%, the held-out documents 82.9%.
  • Yes / no / not stated gains the most: 85.6% whole, against 74.1% with this model's retrieve and 71.3% with solvi-large's. To say "the contract does not say this", the model has to see the whole document.
  • Choice and number questions: retrieve is as good or better (number: 76.9% with retrieve vs 73.8% whole), because the answer sits in one place and BM25 finds it.

Escalation on long documents (2k–8k tokens)

How it was measured.

  • Rows: 3,119 questions at 2k, 4k and 8k tokens (ContractNLI: 1,890 on 81 contracts; held-out documents: 1,229 on 78).
  • Calibration: the act calibrator (logistic regression on the model's confidence features) is fitted on half of the documents and tested on the other half, over 200 random splits.
  • Threshold: a risk-controlled act threshold, P(auto-answer and wrong) ≤ 0.10.
test model accuracy AUROC act answered on its own at risk 0.10 realised risk ECE
ContractNLI solvi-large + retrieve 75.6% 0.784 70.0% 0.098 0.136
ContractNLI this model, full 87.9% 0.826 95.5% 0.101 0.040
held-out documents solvi-large + retrieve 73.3% 0.799 71.6% 0.108 0.109
held-out documents this model, full 83.4% 0.788 85.6% 0.100 0.059

On the held-out documents its act ranking is slightly weaker than solvi-large with retrieve (0.788 vs 0.799; on the 4–8k ones alone 0.756 vs 0.799). It still answers on its own more often at the same risk, because it is right more often.

Short inputs: the solvi-large benchmarks

test solvi-large this model
Fast Decisions dev, zero-shot (macro over domains) 59.4% 59.1%
typed-decisions test, zero-shot 54.5% 54.4%
synthetic states, 9 held-out schemas (all kinds) 97.9% 98.0%
ContractNLI windows, yes / no / not stated (in-distribution) 88.8% 91.4%
ContractNLI windows, evidence quote supports the answer 62.0% 87.3%
Taskmaster-2 held-out slots, span token F1 0.829 0.838
jabr/classifier-benchmark v2, macro 0.694 0.686
decision-models-under-pressure, canonical order, 16 / 64 options 82.2% / 63.1% 81.9% / 62.2%
act AUROC, short inputs: Fast Decisions / typed-decisions / Taskmaster-2 0.762 / 0.668 / 0.743 0.758 / 0.663 / 0.751
act AUROC, short ContractNLI windows 0.854 0.822

The release check it failed

Before training we registered release criteria for a long-input model. One of them compared the act signal with solvi-large's on short test windows: each set's AUROC could drop by at most 0.02. On short ContractNLI windows (≤ 280 words) this model drops by 0.032 (0.854 → 0.822), so it fails that check. The other short-input checks pass (average act AUROC −0.008 against an allowed −0.01; every benchmark within −1.0 point).

The decision. This model is only offered for the opt-in long="full" mode, where texts longer than 512 tokens are read whole. On 2026-09-29, before measuring, we decided that for this mode the short-window act check is replaced by the same check on long documents, where the mode actually runs (the escalation table above; the pass rule: act AUROC at least solvi-large + retrieve − 0.02 on each test, at least as many auto-answers at risk 0.10, realised risk ≤ 0.11). The model passes it. The default model's release rule is unchanged. The failed number is stated here so you can judge it: if your contract inputs are short, use solvi-large.

Speed and the CPU alternative

hardware measurement time
CPU, 4 threads, fp32 (relative to a 512-token pass) whole 2k / 4k / 8k-token text about 5× / 12× / 31×
CPU, same on a typical laptop CPU ≈ 1.6 s per question at 4k, ≈ 4 s at 8k
CPU, same long="retrieve", max_len=2048 (reads at most 2,048 tokens) about 3×
GPU, RTX 3060 Laptop, 6 GB, bf16 (same architecture) 12 questions on 8k-token texts 15.7 s, 2.8 GB peak memory
GPU, RTX PRO 6000 our whole long-document evaluation: 8,274 questions × 4 reading modes 7 min

On a CPU, load the model with a larger retrieve budget instead of reading whole:

m = DecideModel.load("solvi-ai/solvi-large-long", max_len=2048)        # top_k=None: 12 sections of ≈ 170 tokens
part = m.decision("assignment", "...", "contract", type=bool, unknown=True, long="retrieve")

Measured on the same 1,094 questions at 4k–8k tokens, this reads at most 2,048 tokens per question and is as accurate as reading whole: 85.3% vs 84.6% (difference +0.6 points, 95% interval −1.1 to +2.4, bootstrap over documents). Its quotes support the answer a little less often (73.0% vs 76.3%). This budget helps only a model trained on long inputs: solvi-large with max_len=2048 stays at 73%.

Training

  • Initial weights: solvi-ai/solvi-large (Apache-2.0).
  • Recipe: 50 minutes on one GPU (RTX PRO 6000, 96 GB), about 1,300 updates of 128 examples.
    • 15% whole-document examples: a window of 1.5k–8k tokens (log-uniform) containing the evidence, read in one sequence of up to 8,192 tokens.
      • ContractNLI train contracts: 262 contracts, 4,454 hypotheses.
      • CUAD / BillSum / GovReport train documents with open-teacher consensus: 179 documents, 1,485 questions.
    • 85% replay of solvi-large's training mix.
  • Two changes over a plain fine-tune:
    • Evidence spans longer than 40 tokens are clipped to their first 40 tokens in the pointer loss instead of being skipped; the pointer quotes at most 40 tokens. Nearly half of ContractNLI's evidence spans are that long. The same fine-tune without this change quoted supporting evidence on contract windows 53% of the time; with it, 87% (solvi-large: 62%).
    • ContractNLI windows are replayed only for the act head, not in the main loss. This protects the short-input benchmarks.
  • Checkpoint selection: long-document development accuracy (ContractNLI dev, read whole), among checkpoints whose solvi-large development metric did not drop by more than 0.01. Never the tests.
  • Calibration: temperatures and the act calibrator are refitted for this model on development data, not on the tests.

Training data and attribution

Everything in solvi-large's card applies: its data, including the share-alike SQuAD 2.0, BoolQ and VitaminC; this card is their attribution too. Added for long documents:

source license attribution
ContractNLI (train split, whole contracts) CC BY 4.0 Yuta Koreeda, Christopher D. Manning, "ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts", Findings of EMNLP 2021; Hitachi America, Ltd.
CUAD v1 (contracts outside the official test set) CC BY 4.0 The Atticus Project; Hendrycks et al., "CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review", 2021
BillSum (train) CC0-1.0 FiscalNote (Kornilova & Eidelman, 2019); US government works
GovReport (train; the first 2k–7.5k tokens of a report) CC BY 4.0 Huang et al., "Efficient Attentions for Long Document Summarization", NAACL 2021; US GAO / CRS
Questions, labels and quotes for long documents: Qwen/Qwen3-235B-A22B-Instruct-2507, deepseek-ai/DeepSeek-V3.2, openai/gpt-oss-120b Apache-2.0 / MIT / Apache-2.0 (model licenses allow training on outputs) Qwen team (Alibaba Cloud); DeepSeek-AI; OpenAI

The full manifest is LICENSES.json.

Evaluation only, never trained on:

  • ContractNLI test;
  • the 78 held-out long documents;
  • Fast Decisions (dev);
  • typed-decisions test;
  • Taskmaster-2 held-out domains;
  • jabr/classifier-benchmark v2;
  • decision-models-under-pressure.

Limitations

  • A GPU model. Reading 8k tokens whole on a CPU takes seconds per question. solvi warns once per model when long="full" reads more than 2k tokens on a CPU.
  • The act signal on short contract windows is weaker than solvi-large's (AUROC 0.822 vs 0.854); this failed a release check (above). Calibrate on your data.
  • Evaluation limits.
    • ContractNLI results are in-distribution by genre.
    • The 8k bucket is small: 118 questions.
    • The held-out long documents are labelled by teacher-model consensus, not by human annotators.
  • Evidence quotes are at most 40 tokens. A long clause is quoted by its beginning.
  • Short-input benchmarks are within a point of solvi-large but slightly lower on Fast Decisions (−0.3), jabr (−0.8) and decision-models-under-pressure (−0.3 / −0.9).
  • Past 8,192 tokens the model does not see the whole text: solvi retrieves sections within 8,192 tokens, and a question whose evidence BM25 misses can be answered "not stated".
  • Inherited from solvi-large:
    • English only.
    • Not better than GLiNER2.5-Decide on zero-shot choice questions.
    • Several questions per pass is disabled.
    • Not for medical, legal or credit decisions on its own.

Files

  • model.safetensors (bf16).
  • onnx/model_fp16.onnx: full layout, dynamic length: input_ids, attention_mask → logits [B, L, 6].
  • onnx/model_block_fp16.onnx: block layout, for experiments.
  • onnx/parity.json: ONNX vs PyTorch on short inputs and on 2k / 4k-token documents.
  • config.json, tokenizer.json, tokenizer_config.json.
  • solvi_decide.json: capabilities, max_len: 512, max_len_long: 8192, temperatures, act calibrator, multi_question.enabled = false.
  • l14g_format.py, l14f_format.py: the reference input format.
  • LICENSES.json.

Citation

@software{solvi,
  title  = {solvi: verifiable decision systems from catalogs of functions and checks},
  author = {mxkuzn and solvi contributors},
  year   = {2026},
  url    = {https://github.com/solvi-ai/solvi}
}
Downloads last month
11
Safetensors
Model size
0.4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for solvi-ai/solvi-large-long

Quantized
(1)
this model

Datasets used to train solvi-ai/solvi-large-long