decosa-cell-pointer-modernbert-base (v2, value-blind)

Given a paragraph with its numbers tagged and one or more tables, the model says which table cell each number reports: the column for the group the sentence names and the row for the endpoint, statistic, timepoint and population it describes. It can also answer "no cell", and for a computed number (difference, sum, percentage, ratio, relative reduction) it names the operation and its operand cells.

It never sees the cells' values. Tables are written as {column label} under each row label, and "(N=...)" is removed from column labels, so a number can only be placed by what the sentence says, not by where the same value happens to be printed. That is the point: a pointer that sees values tends to cite the cell holding the written (wrong) value and so hides the very error a checker is looking for. Your code then compares the written number with the cell's value.

It was built as the pointer inside a clinical study report (CSR) number checker, where it replaced an LLM call. It runs on a CPU (ONNX Runtime, no GPU needed).

Two models in this repository

Folder Model Use
/ (root) v2, value-blind (the default) Pointing numbers at cells without seeing values
aware/ v1, value-aware twin: same data and schedule as v1, cells written as {Placebo = 52 (25.2)} For comparison and research only. On clinical reports it scores as well as the blind model, but on out-of-domain text it can point at the cell holding the written value (see Results), and it is about twice as slow (longer inputs).

Results

Held-out sets, never trained on. The model is used inside the same CSR number checker end to end (code decides whether a number is wrong); only the pointer changes. The baseline is the same checker with Qwen3.8-27B as the pointer, values shown. All rows use BM25 table retrieval. "Right cell" counts checked numbers whose gold is one cell; "right cell, planted" is the same on the numbers with a planted error (does the pointer still point where the sentence claims when the value is wrong?).

Sets: synthetic test (12 fictional CSRs, 96 planted errors), rewritten test (the same reports with paragraphs rewritten by Qwen3.8, 96 planted errors), ClinicalTrials.gov (30 real phase 3 trials from conditions not in training, their posted results laid out as report tables with template narratives; 178 planted errors).

Set Pointer Errors caught False flags / 100 numbers Right cell Right cell, planted
synthetic test Qwen3.8-27B, values shown 93/96 0.11 99.0% 85.4%
synthetic test this model (v2 blind) 96/96 0.00 100% 100%
synthetic test aware/ (v1 aware) 96/96 0.00 100% 100%
rewritten test Qwen3.8-27B, values shown 93/96 0.00 98.7% 82.1%
rewritten test this model 95/96 0.00 100% 100%
rewritten test aware/ 95/96 0.00 99.9% 98.9%
ClinicalTrials.gov Qwen3.8-27B, values shown 172/178 2.79 97.2% 87.6%
ClinicalTrials.gov this model 176/178 4.08 97.8% 97.7%
ClinicalTrials.gov aware/ 176/178 3.65 98.2% 97.7%

A cascade (this model first, Qwen3.8 with values shown when the model's calibrated confidence is under 0.6, threshold set on dev) did not help: 95/96, 93/96 and 176/178, because the value-aware fallback again cited the cell holding the written value on the numbers it was handed. Use the model alone.

Transfer outside clinical reports is weak. On 10-K MD&A text (two FY2025 annual reports not trained on, 69 figures; gold from a code tie-out that sees values, which favours value-aware pointers):

Pointer Right cell (69)
Qwen3.8-27B, values shown 95.7%
Qwen3.8-27B, values hidden 71.0%
v1 blind 39.1%
v1 aware (aware/) 43.5%, and on a planted swapped-period figure it points at the cell holding the written value
v2 blind (this model) 56.5%

Not ready for financial filings: most misses pick the consolidated line where the sentence names a segment.

Speed: about 0.2-0.45 s of CPU time per checked number on a shared server (ONNX Runtime, 4 threads); a 12-page synthetic report took 25-35 s end to end with no LLM call. Not measured on a dedicated CPU.

Intended use

  • Locating the table cell a number in running text refers to, in clinical study reports and similar results documents (efficacy, safety, baseline and disposition tables), so code can check the number against the cell.
  • As the pointing step of a number-consistency checker, where the separation between "which cell?" (model, value-blind) and "does it match?" (code) is what makes wrong numbers visible.

Not for: deciding by itself whether a number is correct (it points; it does not check), financial filings or other out-of-domain tables without your own evaluation, or any clinical, regulatory or patient-care decision. It is not a validated system and must not be used inside GxP or regulated processes without your own validation and human review.

Limitations (measured, see Results)

  • Trained mostly on clinical report tables; weak on financial segment and note tables (56.5% right cell on 10-K MD&A).
  • The synthetic test reports come from the same generator templates as the training reports (different seeds, drugs, arms and numbers); the rewritten and ClinicalTrials.gov sets are the fairer measures. The ClinicalTrials.gov narratives also use the training template. It has not been tested on a real sponsor CSR.
  • Numbers inside names ("Group 1", "Chondroitin 4&6 Sulfate") get "no cell"; a checker then reports them as untraceable. This is most of the false flags on ClinicalTrials.gov (38 on 30 clean documents).
  • For computed numbers it gives the operands, not their order.
  • One test false-flag fix in the surrounding checker (sums kept within one table) was prompted by errors seen on test, so the test false-flag number is not a fully clean held-out figure; dev showed the same pattern.
  • Latency was measured on a shared machine.
  • model.safetensors (both variants) was rebuilt from the ONNX export after the original training checkpoints were lost; the rebuilt encoder matches ONNX Runtime to a max hidden-state difference of 3e-5. The ONNX files are the ones the results above were measured with.

Training data (every source and its licence)

Source What was used Licence / reuse terms Attribution
Decosa synthetic CSRs Code-generated fictional clinical study reports (fictional drugs, sponsors, trials), half with planted number errors ours (not released) -
Qwen3.8-27B outputs Rewrites of the paragraphs of 140 synthetic reports and 100 ClinicalTrials.gov narratives Model: Apache-2.0 -
ClinicalTrials.gov (API v2) Posted results of 287 phase 3 trials from 50 conditions (257 training, 30 dev), none from the evaluation trials or their conditions; fetched and processed 27 Sep 2026 U.S. NLM terms and conditions (last updated 31 Jan 2023): free reuse with attribution, processing date and a statement of changes Source: ClinicalTrials.gov, U.S. National Library of Medicine. Changes: see NOTICE
TAT-QA (Zhu et al., ACL 2021) Tables and questions from the train split (training) and dev split (model selection) CC BY 4.0 Zhu et al. 2021, github.com/NExTplusplus/TAT-QA
SEC Financial Statement Data Sets, 2026q1 Face statements, with template sentences (v2 stage); the two transfer-test filings excluded U.S. government work, public domain U.S. Securities and Exchange Commission

Not used: FinQA (MIT; its operands are mostly computed, not cell lookups), EMA Policy 0070 clinical data (non-commercial research only). No human labels. The generator, the training data and the evaluation sets are not released.

Held-out discipline: nothing from the evaluation sets (synthetic test seeds, their rewrites, the 30 evaluation trials) or the transfer filings was trained on. Model choice, the temperature and the cascade threshold were set on the model's own dev split.

Model

ModernBERT-base (answerdotai/ModernBERT-base, Apache-2.0) reads [tagged sentences] [SEP] [tables], up to 2,048 tokens. Heads (modeling_cellpointer.py, class Pointer): a pair scorer over (number, cell) plus a "no cell" score, an operation classifier over none / diff / sum / pct / ratio / reduction, and an operand pair scorer. A number is the mean of its tokens and its [#k] tag; a cell is the mean of its {...} tokens. Softmax with temperature 1.4 (in pointer.json) gives calibrated probabilities. v1 was fine-tuned for one epoch on Apple silicon (MPS) on synthetic CSRs, their rewrites, ClinicalTrials.gov tables and TAT-QA; v2 continues v1 with SEC face-statement sentences and a replay of the v1 mix.

Files

  • model.safetensors, config.json, tokenizer.json, tokenizer_config.json, special_tokens_map.json: encoder plus heads (fp32), for PyTorch.
  • onnx/model.onnx (the encoder; output hidden) and onnx/heads.npz (head weights, applied in numpy): the CPU runtime.
  • cellpointer_input.py + csr_tables.py: the exact input rendering. Use them; the model expects this format.
  • modeling_cellpointer.py: the PyTorch module. usage.py: a runnable example.
  • pointer.json: view, max length, temperature, the dev-tuned cascade threshold and the sha256 of every weight file.
  • aware/: the value-aware twin, same layout.
  • eval_summary.json: the numbers above. SHA256SUMS, LICENSE, NOTICE.

Usage

pip install torch transformers safetensors
python usage.py                     # the value-blind model
MODEL_DIR=aware python usage.py     # the value-aware twin
import cellpointer_input as CI
table = CI.table_from_rows("T1", "14.2.1", "Primary endpoint: PASI 75 at Week 16 (full analysis set)", [
    ["", "Placebo (N=103)", "Drug X 40 mg (N=102)"],
    ["Responders, n (%)", "18 (17.5)", "64 (62.7)"],
], header_rows=1)
r = CI.render([{"text": text, "mentions": mentions}], [table], blind=True)
# r.text_a / r.text_b go to the tokenizer as a pair; r.mentions and r.cells give the spans the heads pool over.
# See usage.py for the full forward pass. Output, for the sentence in usage.py:
#   64 [#1] -> T1 R1C2 p=1.00   62.7 [#2] -> T1 R1C2 p=0.99   18 [#3] -> T1 R1C1 p=0.88   17.5 [#4] -> T1 R1C1 p=0.81

Licence

Apache-2.0 for the weights and code in this repository. The base model is Apache-2.0 (ModernBERT, Answer.AI and LightOn). The training data sources above ask for attribution; it is given here and in NOTICE.

Citation and contact

Decosa, decosa.ai. Evaluation: docs/evals/cell-pointer.md in the Decosa API repository (numbers reproduced above).

Downloads last month
-
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for decosaai/decosa-cell-pointer-modernbert-base

Quantized
(76)
this model