YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Synthetic Signed Bank Document Generator

A modular, deterministic, configuration-driven pipeline that takes scanned bank document templates with blank signature fields, fills them with realistic handwritten signatures drawn from a public signature dataset, applies configurable physical/scanner degradation, and exports both the rendered document and a comprehensive per-sample metadata JSON β€” suitable for training and benchmarking document-AI / signature-verification models.

python main.py --config configs/generate.yaml --count 10000

Table of contents

  1. Installation
  2. Project layout
  3. Input assets
  4. Quick start
  5. Configuration
  6. Architecture / pipeline stages
  7. Metadata schema
  8. Verification test cases
  9. Difficulty scoring
  10. Signature box annotation
  11. Determinism & reproducibility
  12. Validation
  13. Visualization
  14. Testing
  15. Extending the pipeline

Installation

Requires Python 3.12+ (developed and tested against 3.11.9 as well; no 3.12-only syntax is used, so 3.11 works too).

python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install -r requirements.txt

Core dependencies: OpenCV, NumPy, SciPy, Pillow, scikit-image, Shapely, PyYAML, Pydantic, Albumentations, OmegaConf, Matplotlib, tqdm, pytest.

Project layout

cc/
β”œβ”€β”€ image/                    # 9 scanned bank document templates (provided)
β”œβ”€β”€ dataset/                  # Public handwritten signature dataset (provided)
β”œβ”€β”€ annotations/              # Signature box annotations, one JSON per template
β”œβ”€β”€ configs/
β”‚   β”œβ”€β”€ default.yaml          # Every config field with its default value
β”‚   β”œβ”€β”€ generate.yaml         # Example large-batch dataset config
β”‚   └── examples/             # Focused example configs (see below)
β”œβ”€β”€ output/                   # Generated PNG + JSON land here
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ config/                # Pydantic schema + OmegaConf loader
β”‚   β”œβ”€β”€ templates/             # TemplateLoader, annotation schema, CV box detector
β”‚   β”œβ”€β”€ signatures/             # SignatureDataset discovery/sampling
β”‚   β”œβ”€β”€ rendering/              # SignaturePreprocessor, Renderer (blending)
β”‚   β”œβ”€β”€ placement/              # PlacementEngine
β”‚   β”œβ”€β”€ augmentation/           # AugmentationEngine + effect implementations
β”‚   β”œβ”€β”€ analysis/                # GeometryAnalyzer, DifficultyAnalyzer, ValidationEngine
β”‚   β”œβ”€β”€ verification/             # Verification test-case registry + CASE_TABLE (case_type 1a-5b)
β”‚   β”œβ”€β”€ metadata/                # MetadataExporter
β”‚   β”œβ”€β”€ visualization/           # Debug figure rendering
β”‚   β”œβ”€β”€ utils/                   # Logging, seeding, geometry, I/O, names
β”‚   └── generator.py             # DatasetGenerator orchestrator (+ multiprocessing)
β”œβ”€β”€ tests/                     # pytest unit + integration tests
β”œβ”€β”€ main.py                    # CLI entry point
└── requirements.txt

Input assets

This repository ships with:

  • image/ β€” 9 real scanned bank document templates (guarantee letters, loan agreements, account mandates), auto-discovered recursively regardless of subfolder layout. Each gets a stable id bank_doc_01 … bank_doc_09, assigned by a deterministic sort of their discovered paths.
  • dataset/ β€” A public handwritten signature dataset (CEDAR-style: genuine/original_<signer>_<sample>.png and forged/forgeries_<signer>_<sample>.png), auto-discovered recursively. Nested dataset/personNNN/*.png-style layouts are also supported β€” the loader falls back to using each signature's parent folder name as its signer group when the CEDAR filename pattern doesn't match, and to an auto-generated id when no identity can be inferred at all. An optional dataset/metadata.csv with path/filename + signer/source_type columns can override both the identity association and genuine/forged labeling for any file, plus an optional script column ("latin" by default) recording the handwriting script a signature is written in. dataset/hindi_bengali/ + the matching rows in dataset/metadata.csv bundle 7 real, named, multi-script signers (Hindi/Devanagari-style, Bengali-style, and Latin cursive) alongside the 55 anonymous CEDAR identities, so generated samples aren't exclusively Latin-script English names β€” SignerIdentityRegistry.display_name() (src/verification/registry.py) returns a metadata.csv-sourced name as-is instead of generating one, and every SignatureRecord/verification metadata block carries its script. These signers have very few reference samples each (1-3, no forged counterpart) β€” records_for_group(..., "forged") falls back to their genuine samples with a logged warning rather than failing, so the 3b skilled-forgery case degrades gracefully rather than erroring for them.

Nothing about either directory's internal layout is hardcoded β€” add more template images or more signature images/folders and they will be picked up automatically on the next run.

Quick start

# 1. Generate 10 samples with all defaults (configs/default.yaml)
python main.py

# 2. Generate a large, varied batch
python main.py --config configs/generate.yaml --count 10000

# 3. Generate one fully reproducible, fixed-identity sample with a debug figure
python main.py --config configs/examples/single_document_demo.yaml --save-visualization

# 4. Override arbitrary nested config values from the command line
python main.py --set placement.rotation_deg.max=6 --set augmentation.coffee_stain.enabled=true

# 5. Restrict to one template, use 8 worker processes, don't overwrite existing files
python main.py --template bank_doc_03 --count 5000 --workers 8

Every run writes output/<prefix>_<NNNNNN>.png + .json pairs, plus output/generation_summary.json (batch statistics) and output/resolved_config.yaml (the fully-resolved config actually used, for provenance/reproduction).

CLI reference

Flag Effect
--config PATH User YAML merged on top of configs/default.yaml
--count N Number of samples
--seed N Master random seed
--workers N Worker process count (1 = sequential, in-process)
--output DIR Output directory
--template ID Force every sample to use one template (e.g. bank_doc_03)
--start-index N First sample number (for resuming/sharding a batch)
--overwrite Regenerate samples even if their files already exist
--save-visualization Also write a debug figure per sample
--set key.path=value Arbitrary dot-path config override (repeatable)
--log-level LEVEL DEBUG / INFO / WARNING / ERROR

Configuration

Configuration is layered: configs/default.yaml (every field, documented) β†’ an optional --config user YAML (merged on top) β†’ --set key=value CLI overrides (highest precedence). The merged result is validated against a Pydantic schema (src/config/schema.py) β€” invalid types, out-of-range values, or unknown keys fail fast with a clear error instead of silently being ignored.

Any placement/augmentation numeric parameter accepts either a fixed scalar or a range, resolved independently per sample by a seeded RNG:

placement:
  scale_factor: 0.84          # fixed
  rotation_deg:
    min: -5
    max: 5                    # sampled uniformly per sample

Every augmentation effect follows the same shape:

augmentation:
  coffee_stain:
    enabled: true
    strength: 0.42             # or {min: ..., max: ...}
    probability: 0.1            # per-sample chance of actually applying
    seed: null                  # optional explicit seed; default derives from (master_seed, sample_index, name)

These effects are covered by 20 configurable families (a couple are strict specializations of another and share one family with a type/tone selector, so the config surface stays coherent):

Family Covers
rotation Extra probabilistic jitter on top of placement.rotation_deg
stroke_dilation Ink stroke thickness variation
motion_blur, gaussian_blur Signature-layer blur
perspective Whole-page perspective warp
scanner_noise, salt_pepper Sensor/impulse noise
jpeg_compression Lossy re-encode artifacts (applied last)
contrast, gamma Photometric shifts
coffee_stain Coffee-ring blotches
stain (type: paper|water|random) Paper stains + water stains
shadow (type: edge|corner|fold|random) Edge / corner / fold shadows
crumple Paper crumpling + wrinkle lines
discoloration (tone: yellow|grey|sepia|random) General discoloration + paper yellowing
page_tilt Crooked scanner feed
dust, scanner_streaks Sensor debris / dirty-glass streaks
ink_fading, ink_bleed Ink density effects

See configs/default.yaml for every field and its default, and configs/generate.yaml / configs/examples/*.yaml for worked examples (a large varied batch, a fixed-identity single-document demo, an "easy" preset, and a "hard/degraded" preset).

Architecture / pipeline stages

Each sample flows through the same 11 stages the spec describes, one class per stage:

# Stage Class
1 Load template TemplateLoader
2 Resolve signature box(es) TemplateLoader (manual annotation, or CV auto-detect fallback)
3 Assign signer name/role src.utils.names + config
4 Select a signature image SignatureDataset
5 Clean/preprocess signature SignaturePreprocessor
6 Fit & position in box PlacementEngine
7 Signature-layer augmentation AugmentationEngine.apply_to_signature
8 Blend onto document Renderer
9 Document-level augmentation AugmentationEngine.apply_to_document
10 Detect rendered signature + geometry GeometryAnalyzer
11 Difficulty score DifficultyAnalyzer
β€” Export MetadataExporter
β€” Verify ValidationEngine

GeneratorContext (in src/generator.py) wires all of the above together for one sample; DatasetGenerator drives it across a batch, either sequentially (workers=1) or via a ProcessPoolExecutor where each worker builds one GeneratorContext (via a pool initializer) and reuses it for every sample it's assigned.

Signature cleaning (SignaturePreprocessor)

Signature dataset images are flat grayscale/RGB scans with no alpha channel. The preprocessor estimates the local paper background level, derives a soft per-pixel alpha from how dark each pixel is relative to that background (gamma-softened to preserve antialiased stroke edges/natural texture rather than a hard binary mask), removes sub-pixel speckle noise via connected-component filtering, and tightly crops to the ink content. A flat ink color (randomly chosen per sample from a small realistic pool β€” black, blue-black, ballpoint blue) is applied to the RGB channels; the alpha channel carries all of the texture.

Placement (PlacementEngine)

Rotates first (about the signature's own center, on an auto-expanding transparent canvas so nothing clips), then scales to fit the box's margin-adjusted interior while preserving aspect ratio (never stretches), then positions via an anchor point (center, top_left, …) plus pixel offsets. A configurable misplacement_probability optionally kicks the signature away from its anchored position by up to misplacement_strength * max(box_w, box_h) in a random direction β€” the mechanism behind intentionally "hard" (badly-placed) samples. A signature is never allowed to shrink below a legible minimum height (16px) even inside a very short box; on the real bundled templates several signature lines are only 12-24px tall, and a naive fit-to-box scale would otherwise render an invisible, sub-pixel signature.

Rendering (Renderer)

Supports multiply, darken, and alpha blend modes, configurable ink opacity, alpha-edge feathering (feather_px), an overall blur to match scan resolution (blur_sigma), a light Gaussian sensor-noise pass baked into every render, and a final mild anti-aliasing blur.

Detection (GeometryAnalyzer)

Detection is scoped to the known signature field (the annotated box, and, if nothing is found there, a widened window around the union of the box and the pre-augmentation placement location) rather than searching the whole page β€” exactly how a real pipeline would use its layout annotations, and what makes the box-vs-detection coverage metrics meaningful in the first place. Ink is separated from the local paper background via adaptive thresholding; a heuristic filter discards thin, wide, solid components (printed "sign here" rule lines) so a blank field is never mistaken for a signature. From the resulting ink mask it computes: axis-aligned bounding box, cv2.minAreaRect minimum rotated rectangle, a convex-hull polygon, centroid, PCA-based orientation, ink pixel count, box_coverage_pct, containment_pct, ink_fill_ratio, polygon_overlap_proxy (IoU), distance from box center, and margin utilization.

Metadata schema

One JSON per sample, always containing the full schema (fields for augmentations that weren't applied are present with a neutral value, not omitted):

{
  "document_id": "sample_000001",
  "template": "bank_doc_03",
  "document_type": "deed_of_guarantee",
  "page": 1, "page_size": [1024, 559], "dpi": 300,
  "signer": {"name": "Jane Doe", "role": "Guarantor (Authorized Signatory)"},
  "signature_source": {"signature_id": "sig_001437", "path": "dataset/genuine/original_14_7.png",
                         "signer_group": "signer_0014", "source_type": "genuine"},
  "placement": {"x_offset": 8.0, "y_offset": -4.0, "rotation_deg": -3.0, "scale_factor": 0.84,
                 "anchor_point": "center", "misplaced": false, ...},
  "rendering": {"blend_mode": "multiply", "ink_opacity": 0.91, "blur_sigma": 0.34, ...},
  "difficulty": {"score": 0.57, "tier": "medium", "misplacement": 0.38, "whitespace": 0.14,
                  "faintness": 0.45, "smallness": 0.31, "low_texture": 0.12},
  "quality": {"degraded": false, "scanner_quality": 0.85, "paper_quality": 0.85, "render_quality": 0.85},
  "verification": {
    "case_type": "3a", "verdict": "REJECT", "reject_code": "IDENTITY_MISMATCH",
    "target_box_id": "sig1", "expected_signer_group": "signer_0014", "expected_signer_script": "latin",
    "actual_signer_group": "signer_0027", "actual_source_type": "genuine", "actual_signer_script": "latin",
    "companion": null
  },
  "augmentation": {
    "rotate_deg": -2, "scale_factor": 0.84, "stroke": 0.12, "noise_amount": 0.06,
    "stain_type": null, "stain_strength": 0.0, "shadow_type": "corner", "shadow_strength": 0.55,
    "coffee_stain_strength": 0.42, "crumple_strength": 0.18, "tilt_deg": 1.3,
    "discoloration_strength": 0.0, "blur_sigma": 0.35, "jpeg_quality": 82, "scanner_noise": 0.02,
    "effects": { "...every single effect, always present, with enabled/applied/strength/probability/seed/params...": {} }
  },
  "geometry": {
    "box_px": [1200, 830, 1530, 950], "target_box": [...], "actual_box": [...],
    "signature_bbox_px": [...], "signature_polygon_px": [[x, y], ...], "convex_hull_px": [...],
    "min_rotated_rect": {"center": [...], "size": [...], "angle_deg": ...},
    "centroid_px": [...], "orientation_deg": ..., "ink_pixels": ..., "detected": true
  },
  "metrics": {
    "box_coverage_pct": 84.2, "containment_pct": 96.3, "ink_fill_ratio": 0.28,
    "polygon_overlap_proxy": 0.71, "distance_from_box_center_px": 12.4, "margin_utilization": 0.6,
    "mean_dark_intensity": 142.1, "gray_stddev": 38.6, "signature_area_ratio": 0.011
  },
  "provenance": {"master_seed": 1234, "sample_index": 1}
}

The augmentation.effects.<name> block is the full audit trail for every one of the 20 effect families: whether it was enabled in config, whether its probability roll actually applied it, the resolved strength, and any effect-specific params (e.g. {"resolved_type": "corner"} for shadow, {"quality": 58} for JPEG compression) β€” this is present even when applied: false, so the schema is identical across every sample regardless of what happened to fire.

quality.degraded is true iff quality.case_type == "1d" or quality.notes contains the substring "low-quality" (case-insensitive) β€” both driven by quality.case_type / quality.notes in config (see configs/examples/hard_degraded.yaml). verification is present (with every field null) on every sample; see the next section for when it's populated.

Verification test cases

Setting quality.case_type to one of 12 recognized codes switches sample generation into scenario-driven mode: instead of a random genuine signature in a random box, the pipeline renders the specific scenario the code describes and stamps the sample's verification metadata block with the matching ground-truth verdict/reject_code β€” producing a labeled benchmark for a downstream verifier. Any other value (or leaving it unset) behaves exactly as before.

Code Scenario Verdict Reject code
1a Correct signer signs normally PASS -
1b Correct signer, natural pen variation (a different genuine sample of the same person) PASS -
1c Multi-signer form, both boxes sign correctly PASS -
1d Correct signer, scan degraded (pair with configs/examples/hard_degraded.yaml) PASS -
2a Box left completely blank REJECT NO_SIGNATURE
2b Printed/typed name only, no handwriting REJECT NO_SIGNATURE
3a Wrong real person signs the box (a different genuine signer) REJECT IDENTITY_MISMATCH
3b Skilled-forgery sample used (CEDAR forged/ sample of the authorized signer) REJECT IDENTITY_MISMATCH
4a Authorized signer, but wrong role's box (the other box's authorized signer signs here instead) REJECT ROLE_MISMATCH
4b Signer from a different form entirely (authorized elsewhere, not on this template) REJECT ROLE_MISMATCH
5a Required signer's box left empty (while a companion box is correctly signed) REJECT REQUIRED_SIGNER_ABSENT
5b Someone else signs in the required signer's place (while a companion box is correctly signed) REJECT REQUIRED_SIGNER_ABSENT
python main.py --config configs/examples/verification_cases.yaml --set quality.case_type=3a --count 10
python scripts/generate_verification_suite.py --per-case 2   # one labeled batch covering all 12 codes

How ground truth is derived

src/verification/registry.py's SignerIdentityRegistry deterministically assigns an "authorized signer" identity (one of the discovered CEDAR signer groups) to every (template_id, box_id) pair, purely as a function of the master seed β€” there's no persisted state, consistent with the rest of the pipeline's seeding model. Every case type is defined declaratively in src/verification/case_types.py (CASE_TABLE) as what should actually be rendered (signature_mode: genuine-correct, blank, printed-text, wrong-signer, skilled-forgery, role-swap, foreign-template-signer, ...) plus whether a companion box on the same page should also be filled with its own correctly-authorized signature (1c/5a/5b β€” the cases whose label only makes sense in a multi-signer-page context). The verification metadata block records both the expected and actual signer group so the label is independently auditable, not just asserted.

Companion-box rendering exists purely for visual/contextual realism (so a "required signer absent" sample actually shows a properly multi-signed page); it does not get its own geometry/difficulty block β€” the schema stays single-target-box-centric, same as it's always been. signature_source is null for the two blank scenarios (2a/5a); the expected identity in that case lives only in verification.expected_signer_group.

Difficulty scoring

Computed exactly as specified, from the post-render, post-augmentation detected geometry (a pure function β€” no randomness β€” so it is always reproducible from its own recorded components):

misplacement = clamp(1 - polygon_overlap_proxy, 0, 1)
whitespace   = clamp(whitespace_ratio, 0, 1)              # whitespace_ratio = 1 - ink_fill_ratio
faintness    = clamp((mean_dark_intensity - 90) / 120, 0, 1)
smallness    = clamp(1 - signature_area_ratio / 0.04, 0, 1)
low_texture  = clamp(1 - gray_stddev / 60, 0, 1)

score = 0.35*misplacement + 0.20*whitespace + 0.20*faintness + 0.15*smallness + 0.10*low_texture

tier = "easy" if score < 0.33 else "medium" if score < 0.66 else "hard"

A note on the bundled templates: smallness uses a fixed reference of 4% of page area. On the 9 real bank forms shipped in image/, the actual signature lines are (realistically) much smaller than that β€” typically 0.2%–1.5% of the page β€” so smallness, and with it the overall score, runs structurally high (mostly medium/hard) even under the "clean" example preset. This is an accurate reflection of how small real bank-form signature fields are relative to a full page, not a bug in the scoring formula, which is intentionally implemented exactly as specified (weights and the 0.04 reference are not something the pipeline silently retunes). If your use case wants a different-looking tier distribution, adjust difficulty.easy_threshold / difficulty.hard_threshold in config, or supply larger signature boxes in your own annotations.

Signature box annotation

Two modes, in order of preference:

  • Mode 2 β€” manual annotation (preferred, used for all 9 bundled templates): annotations/<template_id>.json:

    {
      "id": "bank_doc_03",
      "source_image": "image/guarantee/image.png",
      "document_type": "deed_of_guarantee",
      "pages": [{
        "page": 1, "page_size": [1024, 559], "dpi": 300,
        "signature_boxes": [
          {"id": "sig1", "x1": 595, "y1": 399, "x2": 748, "y2": 431,
           "role": "Guarantor (Authorized Signatory)", "name": "Guarantor"}
        ]
      }]
    }
    

    The 9 bundled annotation files were hand-measured against each template image (using a coordinate-grid overlay for precision) β€” every signature field, its role, and page geometry are captured exactly.

  • Mode 1 β€” automatic CV detection (fallback): if no annotations/<template_id>.json exists for a discovered template, src/templates/box_detector.py looks for long horizontal "sign here" rule lines (morphological opening + contour filtering) and places a candidate box directly above each one. The result is cached back to annotations/<template_id>.json so detection only ever runs once per template β€” subsequent runs (including a fresh checkout with new templates dropped into image/) reuse the cached annotation. If a template has no such lines at all, a generic lower-right fallback box is used so the pipeline never crashes on an unannotated template.

To add a 10th template: drop its image into any subfolder of image/ and either let auto-detection run once, or hand-author its annotation file (fastest way: temporarily overlay a coordinate grid on the image β€” see the _grid_debug pattern used during development β€” and read off box corners).

Determinism & reproducibility

Every random draw anywhere in the pipeline comes from an RNG whose seed is sha256(master_seed | sample_index | stage_name) (src/utils/random_utils.py), never from a shared/global RNG. Consequences:

  • The same (seed, config) always produces byte-identical PNGs and semantically identical JSON, regardless of --workers or scheduling order (verified in tests/test_end_to_end.py, and by hand: a workers=1 run and a workers=4 run of the same seed produce cmp-identical PNGs).
  • Every individual augmentation effect's own seed field in its metadata record is independently derivable/overridable β€” set augmentation.<name>.seed explicitly in config to pin one specific effect while leaving everything else seed-derived.
  • output/resolved_config.yaml captures the exact fully-resolved configuration used for a run, so it can be handed to --config later to reproduce that run exactly (given the same --seed).

Validation

ValidationEngine runs (by default) after every sample and checks: metadata schema completeness, the signature was actually detected, its polygon is non-degenerate and geometrically valid, every bounding box is well-formed, boxes lie on the page (a warning, or an error under validation.strict: true), no NaN/Inf leaked into the JSON, and the difficulty score/tier are exactly reproducible from their own recorded components. Per-sample pass/fail rolls up into output/generation_summary.json. Disable with generation.run_validation: false.

Visualization

--save-visualization writes output/visualizations/<id>_viz.png per sample: the full rendered document with the target box (green), detected bounding box (blue), and detected polygon (red) overlaid; a crop of the signature field; the reconstructed detected-ink mask; a coverage-metrics readout; and a difficulty-component bar chart.

Testing

python -m pytest -q

80 tests covering geometry math (coverage/containment/IoU/margin utilization), the exact difficulty formula (including weight-sum validation and clamping), seed derivation/independence, placement (determinism, aspect-ratio preservation, the minimum-height floor, misplacement gating), signature/template discovery against the real bundled assets (including the multi-script named signer pool), ink detection (including the printed-rule-line false-positive guard), metadata validation, full end-to-end reproducibility, and the case_type-driven verification scenarios (tests/test_verification_cases.py: registry determinism, ground truth for all 12 codes, multi-script signer handling, and backward compatibility with the non-case-type path).

Extending the pipeline

  • New augmentation effect: add a function to src/augmentation/document_effects.py or signature_effects.py, a field to AugmentationsConfig in src/config/schema.py, and a branch in AugmentationEngine.apply_to_document/apply_to_signature.
  • New template: drop an image into image/; annotate manually or let auto-detection + caching handle it.
  • New signature source: drop images anywhere under dataset/ (optionally with a metadata.csv); no code changes needed.
  • Different difficulty calibration: override difficulty.* weights/ thresholds in a config (must still sum to 1.0 β€” enforced by the schema).
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support