YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
- Synthetic Signed Bank Document Generator
Synthetic Signed Bank Document Generator
A modular, deterministic, configuration-driven pipeline that takes scanned bank document templates with blank signature fields, fills them with realistic handwritten signatures drawn from a public signature dataset, applies configurable physical/scanner degradation, and exports both the rendered document and a comprehensive per-sample metadata JSON β suitable for training and benchmarking document-AI / signature-verification models.
python main.py --config configs/generate.yaml --count 10000
Table of contents
- Installation
- Project layout
- Input assets
- Quick start
- Configuration
- Architecture / pipeline stages
- Metadata schema
- Verification test cases
- Difficulty scoring
- Signature box annotation
- Determinism & reproducibility
- Validation
- Visualization
- Testing
- Extending the pipeline
Installation
Requires Python 3.12+ (developed and tested against 3.11.9 as well; no 3.12-only syntax is used, so 3.11 works too).
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
Core dependencies: OpenCV, NumPy, SciPy, Pillow, scikit-image, Shapely, PyYAML, Pydantic, Albumentations, OmegaConf, Matplotlib, tqdm, pytest.
Project layout
cc/
βββ image/ # 9 scanned bank document templates (provided)
βββ dataset/ # Public handwritten signature dataset (provided)
βββ annotations/ # Signature box annotations, one JSON per template
βββ configs/
β βββ default.yaml # Every config field with its default value
β βββ generate.yaml # Example large-batch dataset config
β βββ examples/ # Focused example configs (see below)
βββ output/ # Generated PNG + JSON land here
βββ src/
β βββ config/ # Pydantic schema + OmegaConf loader
β βββ templates/ # TemplateLoader, annotation schema, CV box detector
β βββ signatures/ # SignatureDataset discovery/sampling
β βββ rendering/ # SignaturePreprocessor, Renderer (blending)
β βββ placement/ # PlacementEngine
β βββ augmentation/ # AugmentationEngine + effect implementations
β βββ analysis/ # GeometryAnalyzer, DifficultyAnalyzer, ValidationEngine
β βββ verification/ # Verification test-case registry + CASE_TABLE (case_type 1a-5b)
β βββ metadata/ # MetadataExporter
β βββ visualization/ # Debug figure rendering
β βββ utils/ # Logging, seeding, geometry, I/O, names
β βββ generator.py # DatasetGenerator orchestrator (+ multiprocessing)
βββ tests/ # pytest unit + integration tests
βββ main.py # CLI entry point
βββ requirements.txt
Input assets
This repository ships with:
image/β 9 real scanned bank document templates (guarantee letters, loan agreements, account mandates), auto-discovered recursively regardless of subfolder layout. Each gets a stable idbank_doc_01β¦bank_doc_09, assigned by a deterministic sort of their discovered paths.dataset/β A public handwritten signature dataset (CEDAR-style:genuine/original_<signer>_<sample>.pngandforged/forgeries_<signer>_<sample>.png), auto-discovered recursively. Nesteddataset/personNNN/*.png-style layouts are also supported β the loader falls back to using each signature's parent folder name as its signer group when the CEDAR filename pattern doesn't match, and to an auto-generated id when no identity can be inferred at all. An optionaldataset/metadata.csvwithpath/filename+signer/source_typecolumns can override both the identity association and genuine/forged labeling for any file, plus an optionalscriptcolumn ("latin"by default) recording the handwriting script a signature is written in.dataset/hindi_bengali/+ the matching rows indataset/metadata.csvbundle 7 real, named, multi-script signers (Hindi/Devanagari-style, Bengali-style, and Latin cursive) alongside the 55 anonymous CEDAR identities, so generated samples aren't exclusively Latin-script English names βSignerIdentityRegistry.display_name()(src/verification/registry.py) returns a metadata.csv-sourced name as-is instead of generating one, and everySignatureRecord/verificationmetadata block carries itsscript. These signers have very few reference samples each (1-3, no forged counterpart) βrecords_for_group(..., "forged")falls back to their genuine samples with a logged warning rather than failing, so the3bskilled-forgery case degrades gracefully rather than erroring for them.
Nothing about either directory's internal layout is hardcoded β add more template images or more signature images/folders and they will be picked up automatically on the next run.
Quick start
# 1. Generate 10 samples with all defaults (configs/default.yaml)
python main.py
# 2. Generate a large, varied batch
python main.py --config configs/generate.yaml --count 10000
# 3. Generate one fully reproducible, fixed-identity sample with a debug figure
python main.py --config configs/examples/single_document_demo.yaml --save-visualization
# 4. Override arbitrary nested config values from the command line
python main.py --set placement.rotation_deg.max=6 --set augmentation.coffee_stain.enabled=true
# 5. Restrict to one template, use 8 worker processes, don't overwrite existing files
python main.py --template bank_doc_03 --count 5000 --workers 8
Every run writes output/<prefix>_<NNNNNN>.png + .json pairs, plus
output/generation_summary.json (batch statistics) and
output/resolved_config.yaml (the fully-resolved config actually used, for
provenance/reproduction).
CLI reference
| Flag | Effect |
|---|---|
--config PATH |
User YAML merged on top of configs/default.yaml |
--count N |
Number of samples |
--seed N |
Master random seed |
--workers N |
Worker process count (1 = sequential, in-process) |
--output DIR |
Output directory |
--template ID |
Force every sample to use one template (e.g. bank_doc_03) |
--start-index N |
First sample number (for resuming/sharding a batch) |
--overwrite |
Regenerate samples even if their files already exist |
--save-visualization |
Also write a debug figure per sample |
--set key.path=value |
Arbitrary dot-path config override (repeatable) |
--log-level LEVEL |
DEBUG / INFO / WARNING / ERROR |
Configuration
Configuration is layered: configs/default.yaml (every field, documented) β
an optional --config user YAML (merged on top) β --set key=value CLI
overrides (highest precedence). The merged result is validated against a
Pydantic schema (src/config/schema.py) β invalid types, out-of-range
values, or unknown keys fail fast with a clear error instead of silently
being ignored.
Any placement/augmentation numeric parameter accepts either a fixed scalar or a range, resolved independently per sample by a seeded RNG:
placement:
scale_factor: 0.84 # fixed
rotation_deg:
min: -5
max: 5 # sampled uniformly per sample
Every augmentation effect follows the same shape:
augmentation:
coffee_stain:
enabled: true
strength: 0.42 # or {min: ..., max: ...}
probability: 0.1 # per-sample chance of actually applying
seed: null # optional explicit seed; default derives from (master_seed, sample_index, name)
These effects are covered by 20 configurable families (a couple are
strict specializations of another and share one family with a type/tone
selector, so the config surface stays coherent):
| Family | Covers |
|---|---|
rotation |
Extra probabilistic jitter on top of placement.rotation_deg |
stroke_dilation |
Ink stroke thickness variation |
motion_blur, gaussian_blur |
Signature-layer blur |
perspective |
Whole-page perspective warp |
scanner_noise, salt_pepper |
Sensor/impulse noise |
jpeg_compression |
Lossy re-encode artifacts (applied last) |
contrast, gamma |
Photometric shifts |
coffee_stain |
Coffee-ring blotches |
stain (type: paper|water|random) |
Paper stains + water stains |
shadow (type: edge|corner|fold|random) |
Edge / corner / fold shadows |
crumple |
Paper crumpling + wrinkle lines |
discoloration (tone: yellow|grey|sepia|random) |
General discoloration + paper yellowing |
page_tilt |
Crooked scanner feed |
dust, scanner_streaks |
Sensor debris / dirty-glass streaks |
ink_fading, ink_bleed |
Ink density effects |
See configs/default.yaml for every field and its default, and
configs/generate.yaml / configs/examples/*.yaml for worked examples
(a large varied batch, a fixed-identity single-document demo, an "easy"
preset, and a "hard/degraded" preset).
Architecture / pipeline stages
Each sample flows through the same 11 stages the spec describes, one class per stage:
| # | Stage | Class |
|---|---|---|
| 1 | Load template | TemplateLoader |
| 2 | Resolve signature box(es) | TemplateLoader (manual annotation, or CV auto-detect fallback) |
| 3 | Assign signer name/role | src.utils.names + config |
| 4 | Select a signature image | SignatureDataset |
| 5 | Clean/preprocess signature | SignaturePreprocessor |
| 6 | Fit & position in box | PlacementEngine |
| 7 | Signature-layer augmentation | AugmentationEngine.apply_to_signature |
| 8 | Blend onto document | Renderer |
| 9 | Document-level augmentation | AugmentationEngine.apply_to_document |
| 10 | Detect rendered signature + geometry | GeometryAnalyzer |
| 11 | Difficulty score | DifficultyAnalyzer |
| β | Export | MetadataExporter |
| β | Verify | ValidationEngine |
GeneratorContext (in src/generator.py) wires all of the above together
for one sample; DatasetGenerator drives it across a batch, either
sequentially (workers=1) or via a ProcessPoolExecutor where each worker
builds one GeneratorContext (via a pool initializer) and reuses it for
every sample it's assigned.
Signature cleaning (SignaturePreprocessor)
Signature dataset images are flat grayscale/RGB scans with no alpha channel. The preprocessor estimates the local paper background level, derives a soft per-pixel alpha from how dark each pixel is relative to that background (gamma-softened to preserve antialiased stroke edges/natural texture rather than a hard binary mask), removes sub-pixel speckle noise via connected-component filtering, and tightly crops to the ink content. A flat ink color (randomly chosen per sample from a small realistic pool β black, blue-black, ballpoint blue) is applied to the RGB channels; the alpha channel carries all of the texture.
Placement (PlacementEngine)
Rotates first (about the signature's own center, on an auto-expanding
transparent canvas so nothing clips), then scales to fit the box's
margin-adjusted interior while preserving aspect ratio (never stretches),
then positions via an anchor point (center, top_left, β¦) plus
pixel offsets. A configurable misplacement_probability optionally kicks
the signature away from its anchored position by up to
misplacement_strength * max(box_w, box_h) in a random direction β the
mechanism behind intentionally "hard" (badly-placed) samples.
A signature is never allowed to shrink below a legible minimum height
(16px) even inside a very short box; on the real bundled templates several
signature lines are only 12-24px tall, and a naive fit-to-box scale would
otherwise render an invisible, sub-pixel signature.
Rendering (Renderer)
Supports multiply, darken, and alpha blend modes, configurable ink
opacity, alpha-edge feathering (feather_px), an overall blur to match scan
resolution (blur_sigma), a light Gaussian sensor-noise pass baked into
every render, and a final mild anti-aliasing blur.
Detection (GeometryAnalyzer)
Detection is scoped to the known signature field (the annotated box, and,
if nothing is found there, a widened window around the union of the box and
the pre-augmentation placement location) rather than searching the whole
page β exactly how a real pipeline would use its layout annotations, and
what makes the box-vs-detection coverage metrics meaningful in the first
place. Ink is separated from the local paper background via adaptive
thresholding; a heuristic filter discards thin, wide, solid components
(printed "sign here" rule lines) so a blank field is never mistaken for a
signature. From the resulting ink mask it computes: axis-aligned bounding
box, cv2.minAreaRect minimum rotated rectangle, a convex-hull polygon,
centroid, PCA-based orientation, ink pixel count, box_coverage_pct,
containment_pct, ink_fill_ratio, polygon_overlap_proxy (IoU),
distance from box center, and margin utilization.
Metadata schema
One JSON per sample, always containing the full schema (fields for augmentations that weren't applied are present with a neutral value, not omitted):
{
"document_id": "sample_000001",
"template": "bank_doc_03",
"document_type": "deed_of_guarantee",
"page": 1, "page_size": [1024, 559], "dpi": 300,
"signer": {"name": "Jane Doe", "role": "Guarantor (Authorized Signatory)"},
"signature_source": {"signature_id": "sig_001437", "path": "dataset/genuine/original_14_7.png",
"signer_group": "signer_0014", "source_type": "genuine"},
"placement": {"x_offset": 8.0, "y_offset": -4.0, "rotation_deg": -3.0, "scale_factor": 0.84,
"anchor_point": "center", "misplaced": false, ...},
"rendering": {"blend_mode": "multiply", "ink_opacity": 0.91, "blur_sigma": 0.34, ...},
"difficulty": {"score": 0.57, "tier": "medium", "misplacement": 0.38, "whitespace": 0.14,
"faintness": 0.45, "smallness": 0.31, "low_texture": 0.12},
"quality": {"degraded": false, "scanner_quality": 0.85, "paper_quality": 0.85, "render_quality": 0.85},
"verification": {
"case_type": "3a", "verdict": "REJECT", "reject_code": "IDENTITY_MISMATCH",
"target_box_id": "sig1", "expected_signer_group": "signer_0014", "expected_signer_script": "latin",
"actual_signer_group": "signer_0027", "actual_source_type": "genuine", "actual_signer_script": "latin",
"companion": null
},
"augmentation": {
"rotate_deg": -2, "scale_factor": 0.84, "stroke": 0.12, "noise_amount": 0.06,
"stain_type": null, "stain_strength": 0.0, "shadow_type": "corner", "shadow_strength": 0.55,
"coffee_stain_strength": 0.42, "crumple_strength": 0.18, "tilt_deg": 1.3,
"discoloration_strength": 0.0, "blur_sigma": 0.35, "jpeg_quality": 82, "scanner_noise": 0.02,
"effects": { "...every single effect, always present, with enabled/applied/strength/probability/seed/params...": {} }
},
"geometry": {
"box_px": [1200, 830, 1530, 950], "target_box": [...], "actual_box": [...],
"signature_bbox_px": [...], "signature_polygon_px": [[x, y], ...], "convex_hull_px": [...],
"min_rotated_rect": {"center": [...], "size": [...], "angle_deg": ...},
"centroid_px": [...], "orientation_deg": ..., "ink_pixels": ..., "detected": true
},
"metrics": {
"box_coverage_pct": 84.2, "containment_pct": 96.3, "ink_fill_ratio": 0.28,
"polygon_overlap_proxy": 0.71, "distance_from_box_center_px": 12.4, "margin_utilization": 0.6,
"mean_dark_intensity": 142.1, "gray_stddev": 38.6, "signature_area_ratio": 0.011
},
"provenance": {"master_seed": 1234, "sample_index": 1}
}
The augmentation.effects.<name> block is the full audit trail for every
one of the 20 effect families: whether it was enabled in config, whether
its probability roll actually applied it, the resolved strength, and
any effect-specific params (e.g. {"resolved_type": "corner"} for
shadow, {"quality": 58} for JPEG compression) β this is present even when
applied: false, so the schema is identical across every sample regardless
of what happened to fire.
quality.degraded is true iff quality.case_type == "1d" or
quality.notes contains the substring "low-quality" (case-insensitive) β
both driven by quality.case_type / quality.notes in config (see
configs/examples/hard_degraded.yaml). verification is present (with
every field null) on every sample; see the next section for when it's
populated.
Verification test cases
Setting quality.case_type to one of 12 recognized codes switches sample
generation into scenario-driven mode: instead of a random genuine
signature in a random box, the pipeline renders the specific scenario the
code describes and stamps the sample's verification metadata block with
the matching ground-truth verdict/reject_code β producing a labeled
benchmark for a downstream verifier. Any other value (or leaving it unset)
behaves exactly as before.
| Code | Scenario | Verdict | Reject code |
|---|---|---|---|
1a |
Correct signer signs normally | PASS | - |
1b |
Correct signer, natural pen variation (a different genuine sample of the same person) | PASS | - |
1c |
Multi-signer form, both boxes sign correctly | PASS | - |
1d |
Correct signer, scan degraded (pair with configs/examples/hard_degraded.yaml) |
PASS | - |
2a |
Box left completely blank | REJECT | NO_SIGNATURE |
2b |
Printed/typed name only, no handwriting | REJECT | NO_SIGNATURE |
3a |
Wrong real person signs the box (a different genuine signer) | REJECT | IDENTITY_MISMATCH |
3b |
Skilled-forgery sample used (CEDAR forged/ sample of the authorized signer) |
REJECT | IDENTITY_MISMATCH |
4a |
Authorized signer, but wrong role's box (the other box's authorized signer signs here instead) | REJECT | ROLE_MISMATCH |
4b |
Signer from a different form entirely (authorized elsewhere, not on this template) | REJECT | ROLE_MISMATCH |
5a |
Required signer's box left empty (while a companion box is correctly signed) | REJECT | REQUIRED_SIGNER_ABSENT |
5b |
Someone else signs in the required signer's place (while a companion box is correctly signed) | REJECT | REQUIRED_SIGNER_ABSENT |
python main.py --config configs/examples/verification_cases.yaml --set quality.case_type=3a --count 10
python scripts/generate_verification_suite.py --per-case 2 # one labeled batch covering all 12 codes
How ground truth is derived
src/verification/registry.py's SignerIdentityRegistry deterministically
assigns an "authorized signer" identity (one of the discovered CEDAR signer
groups) to every (template_id, box_id) pair, purely as a function of the
master seed β there's no persisted state, consistent with the rest of the
pipeline's seeding model. Every case type is defined declaratively in
src/verification/case_types.py (CASE_TABLE) as what should actually be
rendered (signature_mode: genuine-correct, blank, printed-text,
wrong-signer, skilled-forgery, role-swap, foreign-template-signer, ...) plus
whether a companion box on the same page should also be filled with its own
correctly-authorized signature (1c/5a/5b β the cases whose label only
makes sense in a multi-signer-page context). The verification metadata
block records both the expected and actual signer group so the label is
independently auditable, not just asserted.
Companion-box rendering exists purely for visual/contextual realism (so a
"required signer absent" sample actually shows a properly multi-signed
page); it does not get its own geometry/difficulty block β the schema
stays single-target-box-centric, same as it's always been. signature_source
is null for the two blank scenarios (2a/5a); the expected identity in
that case lives only in verification.expected_signer_group.
Difficulty scoring
Computed exactly as specified, from the post-render, post-augmentation detected geometry (a pure function β no randomness β so it is always reproducible from its own recorded components):
misplacement = clamp(1 - polygon_overlap_proxy, 0, 1)
whitespace = clamp(whitespace_ratio, 0, 1) # whitespace_ratio = 1 - ink_fill_ratio
faintness = clamp((mean_dark_intensity - 90) / 120, 0, 1)
smallness = clamp(1 - signature_area_ratio / 0.04, 0, 1)
low_texture = clamp(1 - gray_stddev / 60, 0, 1)
score = 0.35*misplacement + 0.20*whitespace + 0.20*faintness + 0.15*smallness + 0.10*low_texture
tier = "easy" if score < 0.33 else "medium" if score < 0.66 else "hard"
A note on the bundled templates: smallness uses a fixed reference of
4% of page area. On the 9 real bank forms shipped in image/, the actual
signature lines are (realistically) much smaller than that β typically
0.2%β1.5% of the page β so smallness, and with it the overall score, runs
structurally high (mostly medium/hard) even under the "clean" example
preset. This is an accurate reflection of how small real bank-form
signature fields are relative to a full page, not a bug in the scoring
formula, which is intentionally implemented exactly as specified (weights
and the 0.04 reference are not something the pipeline silently retunes).
If your use case wants a different-looking tier distribution, adjust
difficulty.easy_threshold / difficulty.hard_threshold in config, or
supply larger signature boxes in your own annotations.
Signature box annotation
Two modes, in order of preference:
Mode 2 β manual annotation (preferred, used for all 9 bundled templates):
annotations/<template_id>.json:{ "id": "bank_doc_03", "source_image": "image/guarantee/image.png", "document_type": "deed_of_guarantee", "pages": [{ "page": 1, "page_size": [1024, 559], "dpi": 300, "signature_boxes": [ {"id": "sig1", "x1": 595, "y1": 399, "x2": 748, "y2": 431, "role": "Guarantor (Authorized Signatory)", "name": "Guarantor"} ] }] }The 9 bundled annotation files were hand-measured against each template image (using a coordinate-grid overlay for precision) β every signature field, its role, and page geometry are captured exactly.
Mode 1 β automatic CV detection (fallback): if no
annotations/<template_id>.jsonexists for a discovered template,src/templates/box_detector.pylooks for long horizontal "sign here" rule lines (morphological opening + contour filtering) and places a candidate box directly above each one. The result is cached back toannotations/<template_id>.jsonso detection only ever runs once per template β subsequent runs (including a fresh checkout with new templates dropped intoimage/) reuse the cached annotation. If a template has no such lines at all, a generic lower-right fallback box is used so the pipeline never crashes on an unannotated template.
To add a 10th template: drop its image into any subfolder of image/ and
either let auto-detection run once, or hand-author its annotation file
(fastest way: temporarily overlay a coordinate grid on the image β see the
_grid_debug pattern used during development β and read off box corners).
Determinism & reproducibility
Every random draw anywhere in the pipeline comes from an RNG whose seed is
sha256(master_seed | sample_index | stage_name) (src/utils/random_utils.py),
never from a shared/global RNG. Consequences:
- The same
(seed, config)always produces byte-identical PNGs and semantically identical JSON, regardless of--workersor scheduling order (verified intests/test_end_to_end.py, and by hand: aworkers=1run and aworkers=4run of the same seed producecmp-identical PNGs). - Every individual augmentation effect's own
seedfield in its metadata record is independently derivable/overridable β setaugmentation.<name>.seedexplicitly in config to pin one specific effect while leaving everything else seed-derived. output/resolved_config.yamlcaptures the exact fully-resolved configuration used for a run, so it can be handed to--configlater to reproduce that run exactly (given the same--seed).
Validation
ValidationEngine runs (by default) after every sample and checks:
metadata schema completeness, the signature was actually detected, its
polygon is non-degenerate and geometrically valid, every bounding box is
well-formed, boxes lie on the page (a warning, or an error under
validation.strict: true), no NaN/Inf leaked into the JSON, and the
difficulty score/tier are exactly reproducible from their own recorded
components. Per-sample pass/fail rolls up into
output/generation_summary.json. Disable with generation.run_validation: false.
Visualization
--save-visualization writes output/visualizations/<id>_viz.png per
sample: the full rendered document with the target box (green), detected
bounding box (blue), and detected polygon (red) overlaid; a crop of the
signature field; the reconstructed detected-ink mask; a coverage-metrics
readout; and a difficulty-component bar chart.
Testing
python -m pytest -q
80 tests covering geometry math (coverage/containment/IoU/margin
utilization), the exact difficulty formula (including weight-sum
validation and clamping), seed derivation/independence, placement
(determinism, aspect-ratio preservation, the minimum-height floor,
misplacement gating), signature/template discovery against the real
bundled assets (including the multi-script named signer pool), ink
detection (including the printed-rule-line false-positive guard), metadata
validation, full end-to-end reproducibility, and the case_type-driven
verification scenarios (tests/test_verification_cases.py: registry
determinism, ground truth for all 12 codes, multi-script signer handling,
and backward compatibility with the non-case-type path).
Extending the pipeline
- New augmentation effect: add a function to
src/augmentation/document_effects.pyorsignature_effects.py, a field toAugmentationsConfiginsrc/config/schema.py, and a branch inAugmentationEngine.apply_to_document/apply_to_signature. - New template: drop an image into
image/; annotate manually or let auto-detection + caching handle it. - New signature source: drop images anywhere under
dataset/(optionally with ametadata.csv); no code changes needed. - Different difficulty calibration: override
difficulty.*weights/ thresholds in a config (must still sum to 1.0 β enforced by the schema).