YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Form Field Detection with Question Extraction

A complete pipeline that extends FFDNet (form field detection) with LayoutLMv3 (text understanding) to programmatically identify each form field, its associated question/prompt, and answer options.

Overview

This project builds on the CommonForms paper (arXiv:2509.16506) by Joe Barrow, which introduced:

  • FFDNet-S (9M params, 72.3 mAP) and FFDNet-L (25M params, 81.0 mAP) for form field detection
  • The CommonForms dataset (480k annotated form pages)

What this pipeline adds

While FFDNet detects where form fields are (text inputs, checkboxes, signatures), it does not tell you what question each field answers. This pipeline:

  1. Detects form fields using FFDNet (or any YOLO-based detector)
  2. Extracts text from the form image via OCR (EasyOCR)
  3. Classifies text tokens as Question/Answer/Header/Other using LayoutLMv3 fine-tuned on FUNSD
  4. Links each field to its associated question text via spatial heuristics
  5. Extracts answer options for choice button fields

Output Format

{
  "num_fields": 3,
  "fields": [
    {
      "field_bbox": [285, 269, 442, 319],
      "field_type": "Text Input",
      "confidence": 0.92,
      "question_text": "Full Name:",
      "question_bbox": [100, 269, 275, 319],
      "answer_options": []
    },
    {
      "field_bbox": [500, 400, 530, 430],
      "field_type": "Choice Button",
      "confidence": 0.88,
      "question_text": "Do you agree?",
      "question_bbox": [300, 400, 490, 430],
      "answer_options": [
        {"text": "Yes", "bbox": [540, 400, 580, 430]},
        {"text": "No", "bbox": [600, 400, 640, 430]}
      ]
    }
  ]
}

Architecture

Input PDF/Image
     |
     v
[EasyOCR] -----> words + bboxes
     |                |
     v                v
[FFDNet-L]     [LayoutLMv3]
     |                |
     v                v
field bboxes     token labels (Q/A/H/O)
     |                |
     +-------> [Spatial Matching]
                    |
                    v
      field -> question link
                    |
                    v
              answer options

Quick Start

Installation

pip install transformers datasets evaluate easyocr Pillow torch numpy opencv-python-headless
pip install ultralytics  # optional, for FFDNet field detection

Usage

from form_understanding_demo import FormUnderstandingPipeline
from PIL import Image

# Initialize pipeline (auto-detects CPU/GPU)
pipeline = FormUnderstandingPipeline(
    layoutlmv3_model="nielsr/layoutlmv3-finetuned-funsd",
    device="cuda",  # or "cpu"
)

# Process an image
image = Image.open("form.png").convert("RGB")

# Option A: Provide pre-detected field boxes from FFDNet
field_boxes = [
    {"bbox": [100, 200, 300, 220], "class_name": "Text Input", "confidence": 0.95},
    {"bbox": [100, 300, 130, 330], "class_name": "Choice Button", "confidence": 0.92},
]
result = pipeline.process_image(image, field_boxes=field_boxes)

# Option B: Without field detection (text classification only)
result = pipeline.process_image(image)

# Print results
for field in result["fields"]:
    print(f"  [{field['field_type']}] Question: '{field['question_text']}'")
    if field["answer_options"]:
        print(f"    Options: {[o['text'] for o in field['answer_options']]}")

CLI Usage

# Process a single image
python form_understanding_demo.py form.png -o result.json

# Process a PDF (requires pdf2image)
python form_understanding_demo.py form.pdf -o result.json

# With pre-detected field boxes (from FFDNet or other detector)
python form_understanding_demo.py form.png --fields detected_boxes.json -o result.json

Models Used

Component Model HF Hub Purpose
Field Detection FFDNet-L jbarrow/FFDNet-L Detect form field boundaries (text input, checkbox, signature)
Field Detection FFDNet-S jbarrow/FFDNet-S Faster variant (5ms/page vs 16ms)
Text Classification LayoutLMv3 nielsr/layoutlmv3-finetuned-funsd Classify tokens as Question/Answer/Header/Other
Base Model LayoutLMv3-base microsoft/layoutlmv3-base LayoutLMv3 base for fine-tuning
OCR EasyOCR N/A Extract word-level text from images

Datasets

  • CommonForms (jbarrow/CommonForms): 480k form pages with field bounding boxes (no text annotations)
  • FUNSD (nielsr/funsd-layoutlmv3): 199 annotated forms with question/answer/header labels

Files

  • form_understanding_demo.py โ€” Complete inference pipeline
  • inference_form_understanding.py โ€” Core pipeline classes (FFDNet + LayoutLMv3 + matching)
  • test_pipeline_demo.py โ€” Test script on FUNSD examples
  • train_layoutlmv3_funsd_v2.py โ€” Fine-tune LayoutLMv3 on FUNSD
  • preprocess_commonforms_for_layoutlmv3.py โ€” Preprocess CommonForms with OCR for pseudo-labeling

How It Works

1. Field Detection (FFDNet)

FFDNet (based on YOLO11) detects three classes of form fields:

  • Text Input (0): Underlines, text boxes
  • Choice Button (1): Checkboxes and radio buttons
  • Signature (2): Signature fields

Trained from scratch on CommonForms at 1216px resolution.

2. Text Classification (LayoutLMv3)

LayoutLMv3 token classification model fine-tuned on FUNSD classifies each word:

  • B-QUESTION / I-QUESTION: The prompt/label text
  • B-ANSWER / I-ANSWER: The content/value text
  • B-HEADER / I-HEADER: Form headers/titles
  • O: Other text

3. Spatial Matching

For each detected field bounding box, the pipeline:

  1. Finds all QUESTION-labeled word groups nearby
  2. Scores candidates by:
    • Horizontal distance (left-of-field, same row is preferred)
    • Vertical distance (above-field, center-aligned is secondary)
  3. Selects the highest-scoring question as the field's prompt
  4. For Choice Button fields, finds nearby text as answer options

Fine-Tuning LayoutLMv3

To improve question detection on your own form dataset:

python train_layoutlmv3_funsd_v2.py

This fine-tunes microsoft/layoutlmv3-base on FUNSD token classification. Key hyperparameters:

  • Learning rate: 5e-5
  • Batch size: 2 (with gradient accumulation 4)
  • Epochs: 20 with early stopping (patience 5)
  • Max length: 512 tokens

Evaluation

The pre-trained LayoutLMv3 (nielsr/layoutlmv3-finetuned-funsd) achieves ~90.8% F1 on FUNSD token classification. Our pipeline on FUNSD test images achieves:

  • Token-level accuracy: 67-97% (varies by document complexity)
  • Question entity detection: 9-26 entities per document
  • Field-question linking: Spatial matching correctly links most fields to their questions

Limitations

  1. OCR dependency: Text extraction depends on EasyOCR quality. Handwritten forms may be challenging.
  2. No multi-page linking: Each page is processed independently.
  3. FUNSD bias: LayoutLMv3 was fine-tuned on FUNSD (English forms); performance may vary on non-English or very different layouts.
  4. Requires pre-detected fields: If FFDNet is not available, the pipeline can still classify text but cannot link to fields.

Future Work

  1. Train LayoutLMv3 on CommonForms-generated pseudo-labels (much larger scale)
  2. Integrate end-to-end training: joint optimization of field detection + question linking
  3. Add multi-page relationship extraction (fields that span pages)
  4. Support for non-English languages using LayoutXLM

Citation

@article{barrow2025commonforms,
  title={CommonForms: A Large, Diverse Dataset for Form Field Detection},
  author={Barrow, Joe},
  journal={arXiv preprint arXiv:2509.16506},
  year={2025}
}

License

The CommonForms dataset and FFDNet models are licensed under Apache 2.0. This pipeline code is released under MIT License.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Paper for blakes/form-field-question-extraction