solvi extract-base
A general field-by-description extractor for solvi: give it a document and a plain description of a field ("the notice period required to terminate the agreement") and it returns the exact span in the text, with a confidence, or says the field is absent. Long documents are read in overlapping 1024-token windows.
It is meant as a starting point: good out of the box on common fields (dates, totals, names), and a better initialization than plain ModernBERT when you fine-tune on your own labeled documents. It is not a reliable zero-shot extractor for arbitrary fields yet — see the numbers below.
Use
from solvi import Answer, Catalog, Question, System
from solvi.extract_long import LongSpanExtractor
ex = LongSpanExtractor.load("solvi-ai/extract-base") # pip install "solvi[model]"
contract_text = open("contract.txt").read()
cat = Catalog()
cat.extract(ex.field("governing_law", "the clause that says which state's or country's law governs the contract"))
@cat.rule("law")
def law(governing_law):
return next((s for s in ["Delaware", "New York", "California"] if s in governing_law), "other")
res = System(cat, [Question("law", "Which law governs?", Answer.choice(["Delaware", "New York", "California", "other"]))]).ask({"doc": contract_text})
print(res["law"].answer, res.computed_state) # the answer and the quoted clause with its offsets
Fine-tune on your documents: ex.fit([(text, description, (start, end) or None), ...], epochs=3), then
ex.tune_threshold(name, held_out) and ex.save(dir).
Architecture
ModernBERT-large encoder + a linear start/end pointer head. Input: [CLS] field description [SEP] document window [SEP].
"No answer" is position 0. The best span over all windows is returned if its score is above the field's threshold (a
default threshold of 0.3 was tuned on held-out data of many fields; tune your own per field with a few labeled documents).
Files: model.safetensors (encoder, bf16), span_head.pt, solvi_extract.json (window settings, thresholds).
Training data
Two epochs over a mix of ~58 000 windows:
- SQuAD 2.0 — 40 000 questions, including unanswerable ones (CC BY-SA 4.0);
- CUAD v1 — 368 training contracts, 33 of the 41 clause types, with CUAD's own clause explanations as descriptions (CC BY 4.0). Held out entirely: Governing Law, Agreement Date, Non-Compete, Exclusivity, Cap On Liability, Effective Date, Uncapped Liability, Competitive Restriction Exception; the 102 official test contracts were never seen;
- CORD v2 — 800 training receipts, 20 field types with written descriptions (CC BY 4.0); tax and change held out.
Evaluation (held-out fields and datasets, one seed)
No labeled examples of the task:
| test | result |
|---|---|
| CORD receipts, fields never trained (tax, change), value match | 46.5% / 67.9% (a model trained on 7 fields only: 14% / 66%) |
| CUAD contracts, 5 typed questions over 5 never-trained clause types | 73.1% (a fine-tuned answerer model trained on 408 contracts: 77.5%) |
| Kleister-NDA (dataset never seen): effective date / jurisdiction / party / term | 43% / 30% / 54% / 59% |
| SROIE receipts (dataset never seen): company / date / address / total | 71% / 95% / 2% / 73% |
With a few labeled documents:
| setup | result |
|---|---|
| CUAD, per-field thresholds and calibration from 40 labeled contracts | 85.5% (vs 77.5% for the answerer trained on 408) |
| SROIE, fine-tuned on 25 / 50 / 100 receipts: field match, this model vs plain ModernBERT-large | 85.8 / 86.1 / 89.2% vs 82.6 / 85.5 / 88.0% |
| SROIE, 100 receipts, typed questions answered by solvi rules | 96.1% vs 94.1% |
Weak spots: addresses (the zero-shot span boundaries do not match), jurisdiction phrased only as a state name, and fields whose wording is far from anything in the training mix. For production, label ~25–100 documents and fine-tune.
Limitations
English only. Trained on receipts, contracts and Wikipedia-style QA; other document types are untested. Scores are
confidences of a span model, not calibrated probabilities — use System.calibrate on held-out documents.
ONNX (browser and CPU)
onnx/model_fp16.onnx — the encoder and span head in one graph (inputs input_ids, attention_mask int64 [batch, length] →
logits float [batch, length, 2]: start and end scores), fp16 weights with float32 inputs and outputs. On 175 test fields it
returns exactly the same spans as the PyTorch model. Runs with onnxruntime (CPU) and onnxruntime-web (WebGPU); windowing and
span decoding follow LongSpanExtractor.predict. Export your own fine-tuned model with tools/export_onnx.py in the solvi repo.
(Dynamic int8 quantization is not provided: it changed half of the spans.)
- Downloads last month
- 29