solvi extract-base

A general field-by-description extractor for solvi: give it a document and a plain description of a field ("the notice period required to terminate the agreement") and it returns the exact span in the text, with a confidence, or says the field is absent. Long documents are read in overlapping 1024-token windows.

It is meant as a starting point: good out of the box on common fields (dates, totals, names), and a better initialization than plain ModernBERT when you fine-tune on your own labeled documents. It is not a reliable zero-shot extractor for arbitrary fields yet — see the numbers below.

Use

from solvi import Answer, Catalog, Question, System
from solvi.extract_long import LongSpanExtractor

ex = LongSpanExtractor.load("solvi-ai/extract-base")          # pip install "solvi[model]"
contract_text = open("contract.txt").read()
cat = Catalog()
cat.extract(ex.field("governing_law", "the clause that says which state's or country's law governs the contract"))

@cat.rule("law")
def law(governing_law):
    return next((s for s in ["Delaware", "New York", "California"] if s in governing_law), "other")

res = System(cat, [Question("law", "Which law governs?", Answer.choice(["Delaware", "New York", "California", "other"]))]).ask({"doc": contract_text})
print(res["law"].answer, res.computed_state)                 # the answer and the quoted clause with its offsets

Fine-tune on your documents: ex.fit([(text, description, (start, end) or None), ...], epochs=3), then ex.tune_threshold(name, held_out) and ex.save(dir).

Architecture

ModernBERT-large encoder + a linear start/end pointer head. Input: [CLS] field description [SEP] document window [SEP]. "No answer" is position 0. The best span over all windows is returned if its score is above the field's threshold (a default threshold of 0.3 was tuned on held-out data of many fields; tune your own per field with a few labeled documents). Files: model.safetensors (encoder, bf16), span_head.pt, solvi_extract.json (window settings, thresholds).

Training data

Two epochs over a mix of ~58 000 windows:

  • SQuAD 2.0 — 40 000 questions, including unanswerable ones (CC BY-SA 4.0);
  • CUAD v1 — 368 training contracts, 33 of the 41 clause types, with CUAD's own clause explanations as descriptions (CC BY 4.0). Held out entirely: Governing Law, Agreement Date, Non-Compete, Exclusivity, Cap On Liability, Effective Date, Uncapped Liability, Competitive Restriction Exception; the 102 official test contracts were never seen;
  • CORD v2 — 800 training receipts, 20 field types with written descriptions (CC BY 4.0); tax and change held out.

Evaluation (held-out fields and datasets, one seed)

No labeled examples of the task:

test result
CORD receipts, fields never trained (tax, change), value match 46.5% / 67.9% (a model trained on 7 fields only: 14% / 66%)
CUAD contracts, 5 typed questions over 5 never-trained clause types 73.1% (a fine-tuned answerer model trained on 408 contracts: 77.5%)
Kleister-NDA (dataset never seen): effective date / jurisdiction / party / term 43% / 30% / 54% / 59%
SROIE receipts (dataset never seen): company / date / address / total 71% / 95% / 2% / 73%

With a few labeled documents:

setup result
CUAD, per-field thresholds and calibration from 40 labeled contracts 85.5% (vs 77.5% for the answerer trained on 408)
SROIE, fine-tuned on 25 / 50 / 100 receipts: field match, this model vs plain ModernBERT-large 85.8 / 86.1 / 89.2% vs 82.6 / 85.5 / 88.0%
SROIE, 100 receipts, typed questions answered by solvi rules 96.1% vs 94.1%

Weak spots: addresses (the zero-shot span boundaries do not match), jurisdiction phrased only as a state name, and fields whose wording is far from anything in the training mix. For production, label ~25–100 documents and fine-tune.

Limitations

English only. Trained on receipts, contracts and Wikipedia-style QA; other document types are untested. Scores are confidences of a span model, not calibrated probabilities — use System.calibrate on held-out documents.

ONNX (browser and CPU)

onnx/model_fp16.onnx — the encoder and span head in one graph (inputs input_ids, attention_mask int64 [batch, length] → logits float [batch, length, 2]: start and end scores), fp16 weights with float32 inputs and outputs. On 175 test fields it returns exactly the same spans as the PyTorch model. Runs with onnxruntime (CPU) and onnxruntime-web (WebGPU); windowing and span decoding follow LongSpanExtractor.predict. Export your own fine-tuned model with tools/export_onnx.py in the solvi repo. (Dynamic int8 quantization is not provided: it changed half of the spans.)

Downloads last month
29
Safetensors
Model size
0.4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for solvi-ai/extract-base

Quantized
(23)
this model
Quantizations
1 model

Datasets used to train solvi-ai/extract-base

Space using solvi-ai/extract-base 1