7b-r64-lr1e-4-nall-s0-e1

Extracts fifteen figures from the primary financial statements of a US SEC 10-K or 10-Q, as JSON. A LoRA fine-tune of allenai/Olmo-3-1025-7B at revision a81bae42db3975be1671e27b9c9a56da1a9f980f, trained on filings with their XBRL facts as the labels. No model graded another model anywhere in producing these numbers.

Results, on filings published after the base model's training cutoff

This model openai/gpt-5.6-luna, zero_shot
Fields correct 97.4% (97.0% to 97.7%) 96.2% (95.8% to 96.6%)
Every field right on a filing 71.9% (68.5% to 75.2%) 58.3% (54.6% to 62.0%)
Answers that are not valid JSON 0.0% (0.0% to 0.0%) 0.0% (0.0% to 0.0%)
Cost per 1,000 extractions US$0.235 served (A40 on_demand at runpod, US$0.490/h, checked 2026-09-26; 1.16 requests a second at 32 in flight, 50% utilisation) US$1.01

705 filings. On filings published before the base model's training cutoff it scores 96.0% (95.7% to 96.4%), against 97.4% (97.0% to 97.7%) after it. That difference is the contamination gap, reported and never used as the headline.

The release gate

Non-inferiority field by field, paired on the same filings, a three-point margin per field and Holm's adjustment across the fifteen, decided by the portfolio's release gate (project 03) on this model's outcomes. A field blocks when the gate cannot rule out a three-point loss on it.

Against Verdict Fields that block Record
openai/gpt-5.6-sol block net_income -2.7 (-4.1 to -1.3), shares_diluted -1.8 (-3.0 to -0.7) c9eb7d588c0cfcf8
openai/gpt-5.6-luna pass none 9acf6e24b0717058

How it fails

Failure Count Per 1,000 fields
Invents a figure the filing does not report 146 13.8
Reads the prior-period column 15 1.4
Drops or misapplies the reporting scale 24 2.3
Reads a loss as a profit, or the reverse 4 0.4
Leaves out a figure the filing does report 55 5.2
Takes a figure from the wrong line 34 3.2
Answer is not valid JSON of the schema 0 0.0

Per field

Field Accuracy (95% CI) Misses by reason
period_end 100.0% (100.0% to 100.0%) none
fiscal_period 100.0% (100.0% to 100.0%) none
revenue 98.7% (97.9% to 99.4%) wrong_value 5, wrong_period 4
cost_of_revenue 91.3% (89.2% to 93.3%) hallucinated 33, missing 28
operating_income 98.6% (97.6% to 99.4%) hallucinated 9, wrong_period 1
net_income 96.0% (94.6% to 97.3%) wrong_value 24, wrong_period 3, scale 1
eps_basic 92.2% (90.1% to 94.0%) hallucinated 46, missing 5, wrong_period 2, wrong_value 1, sign 1
eps_diluted 94.0% (92.2% to 95.7%) hallucinated 29, missing 7, sign 3, wrong_period 2, wrong_value 1
shares_diluted 91.6% (89.5% to 93.6%) hallucinated 29, scale 22, missing 5, wrong_period 3
total_assets 100.0% (100.0% to 100.0%) none
total_liabilities 99.9% (99.6% to 100.0%) missing 1
cash_and_equivalents 99.7% (99.3% to 100.0%) wrong_value 1, missing 1
stockholders_equity 99.9% (99.6% to 100.0%) scale 1
auditor_name 99.7% (99.3% to 100.0%) wrong_value 2
state_of_incorporation 98.9% (98.0% to 99.6%) missing 8

Quantised formats

Each format is judged against this model's bf16 weights over the same filings, and is published only if the lower end of its paired interval is above -1.0% of field accuracy.

Format Paired delta vs bf16 (95% CI) Worst field Published
awq -0.2% (-0.4% to -0.0%) operating_income -0.7% yes
gptq -0.1% (-0.2% to +0.1%) net_income -0.6% yes
gguf-q8_0 -0.0% (-0.1% to +0.0%) shares_diluted -0.4% yes
gguf-q4_k_m -0.1% (-0.3% to +0.0%) net_income -0.6% yes

Training

Base allenai/Olmo-3-1025-7B at a81bae42db3975be1671e27b9c9a56da1a9f980f
Method QLoRA, rank 64, alpha 128, dropout 0.05
Learning rate 0.0001, cosine, warm-up 3%
Examples 5,060 training filings, 713 validation
Epochs 1
Seed 0
Prompt fingerprint e015ec5e057be645
Steps 317
Validation loss 0.01583036407828331
Libraries accelerate 1.15.0, bitsandbytes 0.50.2, peft 0.21.0, python 3.13.15, torch 2.14.0+cu130, transformers 5.17.0, trl 1.13.0

Limitations

  • Trained and measured on US SEC 10-K and 10-Q filings under US GAAP. Nothing here says how it does on IFRS statements, on other regulators' filings, or on documents that are not financial statements.
  • Banks and savings institutions are excluded from the corpus, and funds, trusts and pre-revenue companies are underrepresented. A score here says nothing about their statements. See the dataset's datasheet.
  • The input is the located statements, not a whole filing. Given a whole filing, it has to find them first, which this model was not measured on.
  • Every label is an XBRL fact the company filed, checked to be printed on the page. The label error rate from the hand audit is the ceiling on how far any of these numbers can be trusted.
Downloads last month
3
Safetensors
Model size
7B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Peter-A-P/smallprint-7b

Adapter
(74)
this model
Adapters
1 model