YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

longdoc-kie — paper-inspired document-parsing prototype

Scope honesty first (read this before citing anything here). This repository is a training-free document-parsing prototype that merges techniques from seven papers. It is NOT a faithful reproduction of any of them, and it is not yet the long-document Chinese KIE system the project goal describes. Specifically:

  • Structured KIE: fixed-schema mode implemented, not yet run with a model. longdoc_kie/kie.py does LMDX-style extraction (schema JSON in the prompt, text NNN|XXYY leaf values, explicit null/[] for absent entities, per-line grounding, nested and repeated entities, one owner per entity across overlapping chunks), and run_kie_eval.py scores it on CORD with Donut's micro-F1. It is tested offline only (an oracle and fake models on a real CORD receipt); there are no Qwen numbers yet. Open-schema extraction and relation extraction (key→value links, as in FUNSD/XFUND) are not implemented.
  • Grounding. The parsing pipeline verifies at segment level (formatting-insensitive substring check against the OCR of that segment). KIE values are verified per line (exact inclusion, whitespace runs collapsed) and localized to the OCR words they cover, i.e. below line level; without word boxes the box is interpolated by character position. Sub-line alignment for CJK has not been exercised.
  • Chunking is character-budgeted (~4000 chars, 50% overlap, CPI-style coverage assertion + primary-chunk ownership). Block-aware boundaries (LayIE's self-contained-block rule) are not implemented, and the budget is chars, not model prompt tokens.
  • Chinese long-document KIE is untested. Evaluation ran on llamaindex/ParseBench (single-page, English-first enterprise documents) at the user's direction. Cross-page header→value linking, HUST-CELL/ICDAR-2023-SVRD, XFUND-zh: not exercised. The KIE evaluation target (CORD) is single-page receipts, so long-document KIE is still untested.
  • Metrics are custom diagnostics, implementing ParseBench's published rule semantics (README of the dataset), not the official ParseBench evaluator. Numbers here are not leaderboard-comparable. compare_arms.py re-scores matched documents and reports, per difference, how many documents moved up/down with an exact sign test; bootstrap intervals are reported only from 10 documents. With 4 docs/dimension no difference can reach p < 0.05: 4 of 4 in one direction gives p = 0.125.
  • K=1 greedy Qwen sampling (LMDX uses K=16 with voting) for budget; no voting.
  • CPI window/stride values are defaults, not the paper's (Tan et al. 2026 is closed access; only abstract-level detail was retrievable).

What is here

Technique provenance, per module:

Module Takes from Fix applied per project spec
longdoc_kie/preprocess.py Han et al., ACL 2026 Industry (2026.acl-industry.99) §3.3: edge/contour crop, deskew, bicubic, CLAHE, light denoise Geometry-preserving: crop offsets + deskew angle retained; boxes mapped back to the original page (R^T inverse on 4 corners)
longdoc_kie/surya_stage.py datalab-to/surya-ocr-2 layout+OCR BLOCKIE Creator fix: non-LLM segmentation replaces the LLM Creator
longdoc_kie/chunking.py CPI (Tan et al. 2026, NLPJ 16:100224) + LayIE-LLM (2502.18179) overlapping windows, coverage assertion, primary-chunk ownership
longdoc_kie/lmdx.py LMDX (2309.10952) quantized [x_center,y_center] coords, B=100; segment-ID grounding with substring verification; explicit nulls; formatting-insensitive matching
longdoc_kie/routing.py Multi-VIE (Li et al. 2026, Sci Rep 16:14158) training-free keyword/label router → per-type prompt instructions
longdoc_kie/qwen_client.py Qwen/Qwen3.8-27B via vLLM (thinking disabled) retries with backoff; truncation counted
longdoc_kie/pipeline.py merge + output policy per-segment resolution: primary chunk > overlapping chunk > OCR text; no silent deletion; tables/pictures rendered from Surya; heading levels from Surya
longdoc_kie/score.py ParseBench rule semantics custom diagnostics (see honesty note)
longdoc_kie/kie.py LMDX (2309.10952) §3.2–3.5 fixed-schema extraction: LMDX prompt/completion format, per-line grounding, word-level entity boxes; entities credited to the chunk owning their first line (overlap dedup)
longdoc_kie/kie_score.py Donut evaluator (as used by LMDX on CORD) micro-F1 over (key path, value) pairs; plus whole-item F1, per-field NED, per-field breakdown
longdoc_kie/cord.py CORD (Park et al. 2019) 30-field schema; OCR lines rebuilt by physical row so field boundaries are not leaked; Dataset Viewer loader
compare_arms.py — matched-document comparison; sign test always, bootstrap CIs only from 10 docs
longdoc_kie/htmltext.py — Surya block HTML → text / inline markdown (keeps bold/sup/sub/math)

GEC (Tan et al.) is intentionally replaced by counterfactual evidence-masking tests per the project spec — not yet implemented. SCRI is encoder-training-only and out of scope for this training-free phase.

How to run

pip install numpy opencv-python-headless pillow requests rapidfuzz pytest
python -m pytest tests -q                      # 43 tests (parsing, KIE, stats), no GPU
# full evaluation (GPU, a100-large, vllm/vllm-openai image):
EVAL_BUDGET_S=8700 python3 main_eval.py        # all arms + compare_arms at the end
# single arm:
python3 run_eval.py --config full --per-dim 4 --dims text_content,layout \
    --cache-dir /data/cache --require-cache --qwen-url http://127.0.0.1:8000/v1

Arms: noqwen (Surya only) · noqwen_nopre · full · norouting · nopre · direct (image→Qwen baseline). Pass A builds the shared Surya cache; pass-B arms read it with --require-cache.

KIE on CORD (GPU job, same image and wrapper as main_eval.py; CORD is fetched through the Dataset Viewer API, so nothing beyond requests is needed):

python3 kie_eval_job.py                        # smoke (3 receipts), then both arms on 100
# one arm against a running server:
python3 run_kie_eval.py --qwen-url http://127.0.0.1:8000/v1 --out /data/out/kie_coords
python3 run_kie_eval.py --qwen-url http://127.0.0.1:8000/v1 --no-coords \
    --out /data/out/kie_nocoords               # LMDX Table 5 ablation

Headline results (4 docs/dim shared manifest, matched-doc comparison, ref = noqwen)

Custom diagnostics on 4 documents per dimension. At this size no difference is statistically resolvable: even 4 of 4 documents moving the same way gives sign-test p = 0.125. An earlier version of this table marked differences as "CI excl. 0" using a bootstrap that is too optimistic at n=4; those labels were wrong and are removed.

metric noqwen full direct nopre
word_recall 0.984 0.978 0.969 (−0.016) 0.982
word_f1 0.987 ≈0.985 0.976 ≈0.986
NED (lower=better) 0.028 +0.005 +0.011 ≈0.028
element_pass_rate 0.505 0.505 no layout output 0.523 (+0.019)
block_f1 0.591 0.591 — 0.597
order_tau 0.978 0.978 — 0.978
table_record_f1 0.256 0.256 0.244 0.279
chart_point_pass 0.15 0.15 0.40 0.15
formatting_pass 0.564 0.564 0.714 0.756

Findings, stated plainly:

  1. Qwen verification showed no gain over Surya-only transcription on this English enterprise subset (word_recall slightly lower; nothing resolvable at n=4). This is largely by design: verification only accepts substrings of the OCR. 74 ungrounded predictions across 24 calls fell back to OCR text, so the no-silent-deletion policy held (word_f1 0.987). Logging Qwen's raw outputs would show whether those 74 were real OCR fixes.
  2. Without Han et al. preprocessing, point estimates were slightly higher on formatting, table and layout metrics (not significant at n=4). The preprocessing targets noisy scans, not clean PDFs.
  3. Direct image→Qwen scored lower on text faithfulness (recall, omission, NED) and higher on chart and formatting point estimates. None of these differences is significant at n=4.

Cost/latency (measured, a100-large): Surya-only ≈ 10–18 s/page; full-config Qwen verification ≈ 40 s/page median on this subset; direct ≈ 37 s/page.

Fixed-schema KIE on CORD

Setup follows LMDX §4.1: CORD's own OCR words as input (rebuilt into physical rows, so the model is not handed field boundaries), the 30-field CORD schema, zero-shot, Donut's micro-F1 over (key path, value) pairs. Default dataset is naver-clova-ix/cord-v1, the version LMDX cites; --dataset naver-clova-ix/cord-v2 switches to Donut's corrected release. A receipt whose extraction fails is scored as an empty prediction, not dropped.

Status: implemented and tested offline; no model run yet.

Zero-shot context from LMDX Table 3 (CORD micro-F1): LMDX on PaLM 2-S 66.95 (tuned on non-CORD extraction data first, so not training-free), GPT-4V on the image 64.05, Gemini Pro on OCR 59.57, PaLM 2-S on OCR 55.85, GPT-3.5 on OCR 48.92. The "on OCR" rows (untuned LLMs reading OCR text) are the closest setting to this one. Model, prompt and dataset version still differ, so treat them as context, not a leaderboard.

Known limits: K=1 greedy decoding (LMDX samples 16 and votes); when a value's text occurs twice on its line the first occurrence is used, so its box can point at the wrong copy (the text, and therefore F1, is unaffected); results.json reports any ground-truth field paths the schema cannot produce under gt_keys_outside_schema

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support