decosa-readconf (v2)

A small model that tells a reviewer which cells of a machine-read document to check first. Given a page image and what a document reader (OCR / vision-language reader) returned for it, it scores every table cell and text region with:

  • p_wrong: the probability that the reader's text for that unit is wrong (temperature-calibrated on dev data);
  • the likeliest error type: digit, decimal, merged (merged or split cells), missed_row, handwriting or other.

It does not correct anything and does not read the page itself. It uses 63 features that need no ground truth (the reader's token confidences when available, the layout, column arithmetic such as totals that fail to add up, ink bands in the image, image quality) plus a small character-level encoder over the reader's text. It does not look at the pixels of each cell directly (a variant with a vision encoder was trained but did not do better on real tables and is not released).

It runs on a CPU: the weights are under 1 MB, and a 50-unit page takes about 70 ms including feature extraction.

Results

Baseline: the reader's own confidence, 1 minus the lowest token probability in the unit, which is what a team does today when it sorts by "low confidence". "catch@10" is the share of all wrong units found by reviewing the top 10% of units. Real test tables come from FinTabNet's test shard (26 companies, none of which appear in training); every table was read by PaddleOCR-VL-1.6 and compared with FinTabNet's ground-truth cell text.

Test set Units (wrong) Reader confidence: AUC, catch@10 This model: AUC, catch@10
Real financial tables as published (FinTabNet, 400 tables) 17,873 (979) .868, 56% .906, 69% (AUC +.038, 95% CI +.017 to +.062)
Real tables with a scan/fax/photocopy degradation (the same 400) 74,604 (35,778) .902, 17% .911, 21% (difference not significant; best possible catch@10 is 21%)
Synthetic pages, held-out units 43,598 (5,739) .877, 46% .966, 64%
Synthetic pages, held-out fonts and degradations 27,207 (3,144) .821, 53% .957, 65%

Confidence interval: page-clustered bootstrap, 500 resamples of the 400 test pages. Calibration on real clean tables: ECE .006, Brier .036. The model and training epoch were chosen on a real dev split (FinTabNet train-split tables held out from training); no test split was used for any choice.

By error type on real clean tables (top 10%): merged or split cells 213 of 370 caught (reader confidence 140), decimal errors 22 of 22 (16), digit substitutions 110 of 125 (111).

All numbers, including the other candidates that were not released, are in eval_summary.json.

Intended use

  • Ordering human review of machine-read tables (financial statements, lab tables, forms) so the cells most likely to be misread are checked first, especially merged or split cells, which a reader's token confidence cannot see.
  • As one signal next to the reader's own confidence, not instead of it (see Limitations).

Not for: deciding that a document is correct without review, any clinical, financial or legal decision on its own, or readers other than the one it was trained against without re-labelling. A low p_wrong means none of the patterns it learned were found; it does not mean the reading is right. No decision thresholds ship with the model: rank by p_wrong and review from the top as far as your budget allows, or set thresholds on your own labelled data.

Limitations (measured, see Results)

  • Trained against one reader, PaddleOCR-VL-1.6, and its per-cell token confidences. Other readers make different errors and report confidence differently; expect weaker results until it is re-trained on their output.
  • Digit substitutions: no better than the reader's own confidence on clean tables (110 vs 111 of 125) and much worse on heavily degraded scans (1 vs 1,714 of 2,072 in the top 10%, because merged-cell errors fill the top of the list). If single-digit errors matter most (tie-outs), also flag the reader's lowest-confidence digits, a union of both rankings.
  • Missed rows (measured on synthetic pages only): worse than the reader's confidence (11 vs 39 of 41 in the top 10%).
  • On bad scans p_wrong under-states how often the reader is wrong (ECE .082); use it as a ranking there, not a probability.
  • A first model trained on synthetic pages alone scored below the baseline on real tables (AUC .78-.79 vs .868). It only beat the baseline once real training tables were added, so its reach is limited to what those tables cover: English, US-listed company financial tables. Other layouts, languages and handwriting on real documents are untested.
  • The scanned-table result is dominated by a few pages where the reader loops; the difference from the baseline there is within noise.

Training data (every source and its licence)

Source What was used Licence Attribution
FinTabNet (IBM), train split, via the FinTabNet_OTSL parquet release About 3,000 table images and their ground-truth cell text from 63 companies (1,500 as published, 1,500 with synthetic scan degradations), company-disjoint from the test tables; 85% train, 15% dev CDLA-Permissive-1.0 (no conditions on trained models) X. Zheng et al., "Global Table Extractor (GTE)", WACV 2021; IBM
Synthetic pages generated by a Decosa script 3,000 pages (tables, lab results, forms, handwriting-style values, scan/fax/photocopy/rotation degradations); numbers, labels and names are random Generated by us; not released -
Fonts used to render the synthetic pages DejaVu, Liberation, Noto, Courier Prime, Source Serif 4, Lato, Inconsolata, and handwriting-style Google Fonts (Caveat, Kalam, Patrick Hand, Indie Flower, Homemade Apple, Reenie Beanie, Nanum Pen, Covered By Your Grace, Rock Salt, Dawning of a New Day, Schoolbell) SIL OFL 1.1, Apache-2.0, or the Bitstream Vera / DejaVu licence; no font is included here Each font's authors
PaddleOCR-VL-1.6 (PaddlePaddle) The reader whose output the model scores (its text and token confidences were the training inputs; labels compare them with the ground truth) Apache-2.0 PaddlePaddle authors

No training or evaluation data is redistributed in this repository. The FinTabNet test tables, the synthetic generator and the evaluation sets are not released.

Model

A multilayer perceptron over the 63 standardised features (mean and standard deviation stored as buffers), fused with a 2-layer character CNN over the reader's text (bytes, up to 48), both trained from scratch; heads: one logit for "wrong" and six error-type logits (trained on wrong units only). Temperature 1.16, fitted on the dev split after training. Trained for 4 epochs (4,268 steps, 273,125 units, seed 7). No pretrained base model.

Files: model.safetensors, readconf_config.json (the feature list, error types, temperature), model.py (the network), features.py (turns a page image and a reader page into units and features; numpy and Pillow only), usage.py, eval_summary.json, SHA256SUMS.

Usage

pip install torch safetensors numpy pillow
python usage.py page.png page.json     # prints the 20 units most likely misread

page.json is one page of reader output in this shape (bboxes in the pixels of page.png):

{"elements": [
  {"id": "p1-e2", "kind": "text", "bbox": [110, 139, 1012, 84], "text": "...", "source": "parser",
   "layout_bp": 8930, "conf_bp": 9469, "min_bp": 4423},
  {"id": "p1-e3", "kind": "table", "bbox": [190, 244, 729, 496], "source": "parser", "layout_bp": 7200, "conf_bp": 9877,
   "table": {"rows": 13, "cols": 5, "cells": [
     {"r": 0, "c": 0, "rs": 1, "cs": 1, "text": "Test",
      "tok": {"mean": 9997, "min": 9997, "margin": 9995, "open": 10000, "close": 9999, "n": 1}}]}}]}

Confidences are in basis points (0-10000). Per-cell token confidences (cell["tok"]: mean, min, top-1 vs top-2 margin, first and last token, token count) are optional but help a lot; see features.py for every field it reads.

import json, numpy as np, torch
from PIL import Image
from safetensors.torch import load_file
import features as F
from model import TYPES, ReadConf, encode_text

cfg = json.load(open("readconf_config.json"))
m = ReadConf(cfg["n_feat"], vision_layers=0); m.load_state_dict(load_file("model.safetensors")); m.eval()
units = F.page_units(Image.open("page.png").convert("RGB"), json.load(open("page.json")))
with torch.no_grad():
    logit, type_logits = m(torch.tensor(np.array([u["x"] for u in units], np.float32)), encode_text([u["text"] for u in units]))
p = m.p_wrong(logit)
for i in p.argsort(descending=True)[:10].tolist():
    print(units[i]["uid"], units[i]["text"], round(float(p[i]), 3), TYPES[int(type_logits[i].argmax())])

Licence

Apache-2.0 for the weights and code in this repository. The training data sources above are credited here and in NOTICE.

Contact

Decosa, decosa.ai.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
214k params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support