YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Form Field Detection with Question Extraction
A complete pipeline that extends FFDNet (form field detection) with LayoutLMv3 (text understanding) to programmatically identify each form field, its associated question/prompt, and answer options.
Overview
This project builds on the CommonForms paper (arXiv:2509.16506) by Joe Barrow, which introduced:
- FFDNet-S (9M params, 72.3 mAP) and FFDNet-L (25M params, 81.0 mAP) for form field detection
- The CommonForms dataset (480k annotated form pages)
What this pipeline adds
While FFDNet detects where form fields are (text inputs, checkboxes, signatures), it does not tell you what question each field answers. This pipeline:
- Detects form fields using FFDNet (or any YOLO-based detector)
- Extracts text from the form image via OCR (EasyOCR)
- Classifies text tokens as Question/Answer/Header/Other using LayoutLMv3 fine-tuned on FUNSD
- Links each field to its associated question text via spatial heuristics
- Extracts answer options for choice button fields
Output Format
{
"num_fields": 3,
"fields": [
{
"field_bbox": [285, 269, 442, 319],
"field_type": "Text Input",
"confidence": 0.92,
"question_text": "Full Name:",
"question_bbox": [100, 269, 275, 319],
"answer_options": []
},
{
"field_bbox": [500, 400, 530, 430],
"field_type": "Choice Button",
"confidence": 0.88,
"question_text": "Do you agree?",
"question_bbox": [300, 400, 490, 430],
"answer_options": [
{"text": "Yes", "bbox": [540, 400, 580, 430]},
{"text": "No", "bbox": [600, 400, 640, 430]}
]
}
]
}
Architecture
Input PDF/Image
|
v
[EasyOCR] -----> words + bboxes
| |
v v
[FFDNet-L] [LayoutLMv3]
| |
v v
field bboxes token labels (Q/A/H/O)
| |
+-------> [Spatial Matching]
|
v
field -> question link
|
v
answer options
Quick Start
Installation
pip install transformers datasets evaluate easyocr Pillow torch numpy opencv-python-headless
pip install ultralytics # optional, for FFDNet field detection
Usage
from form_understanding_demo import FormUnderstandingPipeline
from PIL import Image
# Initialize pipeline (auto-detects CPU/GPU)
pipeline = FormUnderstandingPipeline(
layoutlmv3_model="nielsr/layoutlmv3-finetuned-funsd",
device="cuda", # or "cpu"
)
# Process an image
image = Image.open("form.png").convert("RGB")
# Option A: Provide pre-detected field boxes from FFDNet
field_boxes = [
{"bbox": [100, 200, 300, 220], "class_name": "Text Input", "confidence": 0.95},
{"bbox": [100, 300, 130, 330], "class_name": "Choice Button", "confidence": 0.92},
]
result = pipeline.process_image(image, field_boxes=field_boxes)
# Option B: Without field detection (text classification only)
result = pipeline.process_image(image)
# Print results
for field in result["fields"]:
print(f" [{field['field_type']}] Question: '{field['question_text']}'")
if field["answer_options"]:
print(f" Options: {[o['text'] for o in field['answer_options']]}")
CLI Usage
# Process a single image
python form_understanding_demo.py form.png -o result.json
# Process a PDF (requires pdf2image)
python form_understanding_demo.py form.pdf -o result.json
# With pre-detected field boxes (from FFDNet or other detector)
python form_understanding_demo.py form.png --fields detected_boxes.json -o result.json
Models Used
| Component | Model | HF Hub | Purpose |
|---|---|---|---|
| Field Detection | FFDNet-L | jbarrow/FFDNet-L | Detect form field boundaries (text input, checkbox, signature) |
| Field Detection | FFDNet-S | jbarrow/FFDNet-S | Faster variant (5ms/page vs 16ms) |
| Text Classification | LayoutLMv3 | nielsr/layoutlmv3-finetuned-funsd | Classify tokens as Question/Answer/Header/Other |
| Base Model | LayoutLMv3-base | microsoft/layoutlmv3-base | LayoutLMv3 base for fine-tuning |
| OCR | EasyOCR | N/A | Extract word-level text from images |
Datasets
- CommonForms (
jbarrow/CommonForms): 480k form pages with field bounding boxes (no text annotations) - FUNSD (
nielsr/funsd-layoutlmv3): 199 annotated forms with question/answer/header labels
Files
form_understanding_demo.pyโ Complete inference pipelineinference_form_understanding.pyโ Core pipeline classes (FFDNet + LayoutLMv3 + matching)test_pipeline_demo.pyโ Test script on FUNSD examplestrain_layoutlmv3_funsd_v2.pyโ Fine-tune LayoutLMv3 on FUNSDpreprocess_commonforms_for_layoutlmv3.pyโ Preprocess CommonForms with OCR for pseudo-labeling
How It Works
1. Field Detection (FFDNet)
FFDNet (based on YOLO11) detects three classes of form fields:
- Text Input (0): Underlines, text boxes
- Choice Button (1): Checkboxes and radio buttons
- Signature (2): Signature fields
Trained from scratch on CommonForms at 1216px resolution.
2. Text Classification (LayoutLMv3)
LayoutLMv3 token classification model fine-tuned on FUNSD classifies each word:
- B-QUESTION / I-QUESTION: The prompt/label text
- B-ANSWER / I-ANSWER: The content/value text
- B-HEADER / I-HEADER: Form headers/titles
- O: Other text
3. Spatial Matching
For each detected field bounding box, the pipeline:
- Finds all QUESTION-labeled word groups nearby
- Scores candidates by:
- Horizontal distance (left-of-field, same row is preferred)
- Vertical distance (above-field, center-aligned is secondary)
- Selects the highest-scoring question as the field's prompt
- For Choice Button fields, finds nearby text as answer options
Fine-Tuning LayoutLMv3
To improve question detection on your own form dataset:
python train_layoutlmv3_funsd_v2.py
This fine-tunes microsoft/layoutlmv3-base on FUNSD token classification. Key hyperparameters:
- Learning rate: 5e-5
- Batch size: 2 (with gradient accumulation 4)
- Epochs: 20 with early stopping (patience 5)
- Max length: 512 tokens
Evaluation
The pre-trained LayoutLMv3 (nielsr/layoutlmv3-finetuned-funsd) achieves ~90.8% F1 on FUNSD token classification. Our pipeline on FUNSD test images achieves:
- Token-level accuracy: 67-97% (varies by document complexity)
- Question entity detection: 9-26 entities per document
- Field-question linking: Spatial matching correctly links most fields to their questions
Limitations
- OCR dependency: Text extraction depends on EasyOCR quality. Handwritten forms may be challenging.
- No multi-page linking: Each page is processed independently.
- FUNSD bias: LayoutLMv3 was fine-tuned on FUNSD (English forms); performance may vary on non-English or very different layouts.
- Requires pre-detected fields: If FFDNet is not available, the pipeline can still classify text but cannot link to fields.
Future Work
- Train LayoutLMv3 on CommonForms-generated pseudo-labels (much larger scale)
- Integrate end-to-end training: joint optimization of field detection + question linking
- Add multi-page relationship extraction (fields that span pages)
- Support for non-English languages using LayoutXLM
Citation
@article{barrow2025commonforms,
title={CommonForms: A Large, Diverse Dataset for Form Field Detection},
author={Barrow, Joe},
journal={arXiv preprint arXiv:2509.16506},
year={2025}
}
License
The CommonForms dataset and FFDNet models are licensed under Apache 2.0. This pipeline code is released under MIT License.