YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Document PAD (bona fide vs screen) - experiment ledger

Mirror of experiments/ledger.md in the RunPod workspace; pushed to the root of US10F/document-pad-runpod as README.md. One row per experiment. Protocol, gates and backlog: see CLAUDE.md in this repo.

Gates (owner, 2026-09-24, corrected; on each gating held-out set - DLC and SaudiDocs; KID = stress test): AUC >= 0.88, APCER <= 0.10, BPCER <= 0.10, accuracy >= 0.85; soft target EER <= 0.15 (ideally <= 0.10). Threshold chosen on one subject-disjoint part of the held-out set, reported on the rest. Earlier rows were judged against the old gates (AUC .90, ACER/APCER/BPCER .08).

Legacy runs (Colab, before this ledger) - for reference, not re-run

All EfficientNetV2-B0 @ 896x576 unless noted, screen-only, val = DLC 5000/5000 (80 docs), threshold 0.5. All block-4 runs used the identical DLC val / KID 5k images (exp_001: the results.json fingerprints hashed cache paths; the printed ones match). Epoch-to-epoch swing of held-out AUC is ~0.03 (exp_006), so single-run differences under ~0.02 are not meaningful.

id train variable best DLC AUC (epoch) acc@0.5 recall@0.5
L1 block1_2 KID 5k/5k no SaudiDocs, no aug 0.717 (3) 0.559 0.985
L2 block4_1 KID 5k/5k + Saudi uncropped + SaudiDocs 0.876 (3) 0.783 0.936
L3 block4_1 KID 5k/5k + Saudi cropped cropped Saudi 0.891 (3) 0.764 0.941
L4 block4_1 KID full + Saudi cropped full KID 0.910 (2) 0.680 0.971
L5 block4_2 KID 5k/5k + Saudi uncropped + augmentation 0.880 (2) 0.806 0.943
L6 block4_2 KID 5k/5k + Saudi cropped + augmentation 0.929 (2) 0.757 0.973
L7 block4_2 as L6, unfreeze 1, lr_bb 1e-5 regularisation 0.922 (2) 0.695 0.970
L8 block4_2 KID full + Saudi cropped + aug full KID 0.911 (3) 0.692 0.998
L9 block10 ConvNeXt-T 224x224, TPO->KID full+Saudi backbone+size+pretrain 0.854 (2), EER 0.224 0.633 0.990

Pattern in every run: held-out AUC peaks at epoch 2-3 and then falls while train AUC -> 1.0 (source-domain overfit); recall 0.94-0.998, BPCER 0.33-0.72 at threshold 0.5.

Experiments on this pod

id hypothesis single variable baseline headline (worst held-out) keep? lesson
exp_001_threshold_calibration 0.83 plateau is a threshold problem threshold rule (post-hoc, no training) L1-L9 best L6: AUC 0.929, EER 0.127; @0.5 ACER 0.243 (BPCER 0.460) -> cross-fold EER thr 0.893: ACER 0.133, BPCER 0.128, APCER 0.137, acc 0.868; oracle min ACER 0.126 keep (calibrate every run) calibration fixes the accuracy gate but no threshold reaches ACER<=0.08: min ACER ~= EER, corr(AUC,EER)=-0.98, so the remaining gap is ranking, not the cut
exp_002_baseline_valDLC (planned: pod reproduces L6) accidental: HEAD_INIT timm - timm's 1-output head init is uniform +-1 (std 0.58, ~35x the legacy nn.Linear init); otherwise L6 config, held-out DLC. Its results.json says 'none' - read this row instead L6 DLC best 0.948 (ep 8/10), last 0.941, EER 0.114; cf-EER thr 0.880: ACER 0.114, BPCER 0.115, APCER 0.113, acc 0.886 promising (vs exp_006: +0.027 best, +0.041 last AUC; higher every epoch >=2) - confirming in exp_012/013 beats legacy L6 (0.929/EER 0.127) - but judge only against exp_006 + seed noise; held-out AUC swings 0.908-0.948 across epochs
exp_006_baseline_valDLC pod reproduces L6 none (L6 config, legacy head init, batch 16) - held-out DLC L6 DLC best 0.921 (ep 5/10), last 0.900, EER 0.157; cf-EER thr 0.946: ACER 0.156, BPCER 0.156, APCER 0.156, acc 0.844 reference reproduces L6 (0.929) within the epoch-to-epoch swing (0.880-0.921); train loss ep1 0.397 = L6 0.396
exp_012_headinit_timm_valKID timm head init gain holds on KID fold HEAD_INIT timm exp_003 KID best 0.686, last 0.682, mean 0.675 (base 0.676), EER 0.363 discard for KID (DLC gain awaits exp_013) the exp_002 DLC gain does not transfer to the KID fold
exp_013_headinit_timm_valDLC_seed7 timm head init gain is not a seed accident HEAD_INIT timm, SEED 7 exp_005 superseded (EffNet baseline), not run - ConvNeXt replaced the baseline before this ran
exp_003_baseline_valKID second LODO fold train DLC 5k/5k + Saudi, held-out KID34K exp_006 KID best 0.690 (ep 8/10), last 0.686, EER 0.353; cf-EER thr 0.788: ACER 0.355, acc 0.645 reference - this is now the headline (worst) fold KID is the hard fold. Train AUC 1.000 on DLC+Saudi: shortcut learning; see diagnosis below
exp_014_landscape_valKID removing the orientation cue helps KID ORIENT landscape exp_003 KID best 0.713 (ep 3), last 0.655, mean 0.664 (base 0.676), EER 0.345; cf ACER 0.346 discard best epoch is noise; mean/last worse. Orientation is real in the data but not what limits KID. BPCER@0.5 jumped to 0.80 (KID scored as attack)
exp_015_landscape_valDLC ... and does not hurt DLC ORIENT landscape exp_006 cancelled - not worth running after exp_014
exp_016_patch384_valKID local artefacts transfer PATCH 384 exp_003 KID best 0.690 (ep 7), last 0.669, mean 0.664, EER 0.363 discard patches do not escape the KID ceiling either
exp_017_convnext_tiny_valKID capacity helps KID convnext_tiny (in12k_ft_in1k) exp_003 KID best 0.746 (ep 7), last 0.732, mean 0.726 (base 0.676), EER 0.313; cf ACER 0.311 keep - first real KID gain backbone matters for the stress fold; above baseline at every epoch >= 2; 10 min/epoch
exp_025_convnext_dinov3_valKID self-supervised web-scale pretraining transfers convnext_tiny.dinov3_lvd1689m exp_017 KID best 0.628 (ep 10), mean 0.548 (exp_017 0.726), EER 0.420 discard DINOv3 weights much worse than in12k with the same recipe; exp_029 (DLC) moved to the end of the queue
exp_026_convnext_small_valKID more capacity convnext_small.in12k_ft_in1k exp_017 queued - -
exp_004_baseline_valSaudi deployment fold (smoke) train KID 5k/5k + DLC 5k/5k, held-out SaudiDocs exp_006 Saudi best 0.871 (ep 8), last 0.867, EER 0.206; @0.5 APCER 0.226 BPCER 0.182 (2 subjects: no cross-fold) reference below the SaudiDocs 0.90 gate without Saudi in training
exp_018_sauditight_valKID Saudi background is a shortcut SAUDI_TIGHT (crop to report card box) exp_003 KID best 0.674 (ep 10), mean 0.665, EER 0.369 discard background is not what limits KID
exp_020_unfreeze1_valKID freezing more limits source overfit UNFREEZE 4 -> 1 exp_003 superseded (EffNet baseline), not run - ConvNeXt replaced the baseline before this ran
exp_022_gray_valDLC colour cast is the cue GRAYSCALE True exp_006 DLC best 0.933 (ep 5), last 0.877, mean 0.894 (base 0.906), EER 0.140 discard removing colour does not fix the tinted-bona errors; colour is not the only cue
exp_023_gray_valKID colour cast is the cue GRAYSCALE True exp_003 KID best 0.642, mean 0.628 (base 0.676), EER 0.400 discard grayscale hurts the stress fold too; colour carries real signal
exp_024_wbaug_valDLC colour made unreliable AUG p_wb 0.8 gains 0.6-1.4 + p_gray 0.2 exp_006 superseded (EffNet baseline), not run - ConvNeXt replaced the baseline before this ran
exp_021_dinov2_frozen_valKID frozen foundation features transfer frozen DINOv2 ViT-S/14 + linear head (1,153 trainable) exp_003 KID best 0.586 (ep 1), mean 0.581, EER 0.449; train loss stuck ~0.35 discard frozen semantic features do not carry recapture evidence (cannot even fit train); the signal is low-level texture that needs fine-tuning
exp_019_sauditight_valDLC ... no harm on DLC SAUDI_TIGHT exp_006 superseded (EffNet baseline), not run - ConvNeXt replaced the baseline before this ran
exp_005_baseline_valDLC_seed7 seed noise << effects SEED 42 -> 7 (same images) exp_006 superseded (EffNet baseline), not run - ConvNeXt replaced the baseline before this ran
exp_007_posweight_valDLC balancing the loss lowers BPCER@0.5 POS_WEIGHT balanced (0.84) exp_006 superseded (EffNet baseline), not run - ConvNeXt replaced the baseline before this ran
exp_008_convnext_tiny_valDLC capacity + sweep-best family at full res backbone -> convnext_tiny (in12k_ft_in1k) @896x576, all 4 stages trainable exp_006 DLC best 0.973 (ep 8/10), last 0.963, mean 0.955 (base 0.906), EER 0.082; cf-EER thr 0.981: APCER 0.083, BPCER 0.081, ACER 0.082, acc 0.918 keep - NEW BASELINE within 0.003 of every DLC gate; AUC/acc pass. Epoch swing 0.927-0.973. 10 min/epoch
exp_027_convnext_valSaudi ConvNeXt lifts the deployment fold convnext_tiny, Saudi held out exp_004 Saudi best 0.933 (ep 4), last 0.922, EER 0.149; @0.5 APCER 0.221 BPCER 0.070; per subject AUC fahad 0.957 / yousef 0.936 (median attack P 0.68 vs 0.97) keep - Saudi AUC gate passed (EffNet 0.871) error-rate gates far on Saudi (EER 0.15); fahad's screen captures are the weak spot
exp_028_convnext_ema_valDLC EMA smooths the swing EMA 0.999 on convnext exp_008 DLC best 0.980 (ep 5), last 0.979, mean 0.969, EER 0.075; cf-EER: ACER 0.0755, BPCER 0.074, APCER 0.078, acc 0.924 - PASSES new gates keep - best DLC model EMA removes the epoch swing (last ~ best) and lowers EER 0.082 -> 0.075
exp_033_convnext_ema_valSaudi EMA helps the binding Saudi fold EMA 0.999, Saudi held out exp_027 Saudi best 0.933 (ep 4), last 0.915, EER 0.145; data thr 0.074: APCER 0.144 BPCER 0.145 acc 0.855; leave-one-subject-out: APCER 0.170 BPCER 0.139 acc 0.811 - FAIL discard (no change vs exp_027) EMA does not help the Saudi fold; the Saudi gap is not an epoch-noise problem
exp_034_indomain_valKID 1% production gate reachable in a seen domain PROTOCOL exception: KID split by user (32 train / 14 val), all 3 datasets in training, recipe convnext + colour aug + EMA exp_028 queued - -
exp_035_indomain_valDLC same, DLC PROTOCOL exception: DLC split by document (57 / 23), recipe convnext + colour aug + EMA exp_037 recipe queued - -
exp_029_convnext_dinov3_valDLC SSL pretraining transfers convnext_tiny.dinov3_lvd1689m exp_008 queued - -
exp_030_convnext_small_valDLC more capacity convnext_small (+EMA) exp_028 queued - -
exp_031_convnext_valDLC_seed7 not a seed accident SEED 7 (+EMA) exp_028 queued - -
exp_032_convnext_wbaug_valDLC colour made unreliable (not removed) AUG p_wb 0.8 gains 0.6-1.4 + p_gray 0.2 exp_008 DLC best 0.988 (ep 7), last 0.974, mean 0.979, EER 0.050; at data thr 0.972: APCER 0.050, BPCER 0.050, acc 0.950; cf APCER 0.050 BPCER 0.051 - PASS (EER <= 0.10) keep - best DLC model confirms the colour-cast shortcut: making colour unreliable (not removing it, cf. exp_022) cuts EER 0.082 -> 0.050
exp_036_convnext_wbaug_valSaudi colour shortcut also explains the Saudi failure colour-cast aug, Saudi held out exp_027 Saudi best 0.911 (ep 9), EER 0.171; data thr: APCER/BPCER 0.171, acc 0.829; LOSO APCER 0.247 BPCER 0.146 - FAIL discard colour aug helps DLC but hurts Saudi (0.933 -> 0.911): the Saudi gap is not colour
exp_038_convnext_sauditight_valSaudi the model reads the card's surroundings SAUDI_TIGHT on the Saudi held-out set exp_027 queued - -
exp_037_convnext_wbaug_ema_valDLC EMA + colour aug add up EMA 0.999 on exp_032 recipe exp_032 queued - -
exp_009_effv2s_valDLC capacity within family backbone -> tf_efficientnetv2_s exp_006 superseded (EffNet baseline), not run - ConvNeXt replaced the baseline before this ran
exp_010_ema_valDLC weight averaging smooths the early, noisy held-out peak EMA_DECAY 0.999 exp_006 superseded (EffNet baseline), not run - ConvNeXt replaced the baseline before this ran
exp_011_patch384_valDLC local artefacts transfer, layout does not input: random 384 crops, eval mean of 3x2 grid exp_006 superseded (EffNet baseline), not run - ConvNeXt replaced the baseline before this ran

exp_001 details

  • Re-scored 9 legacy checkpoints on the historical DLC val draw (10k, 80 docs). Block-4 AUCs reproduce to 4 decimals (0.8759 ... 0.9286), so the pod preprocessing is exact.
  • Fingerprint "groups" were an artifact: the printed split md5s (over source paths) are identical in every legacy block-4 run - DLC val bona b800946b1e7f / attack 2bafc288edf2, KID 5k a116f90d8132/f9abf358b93b - and match the pod. The differing results.json values hashed the cache paths (md5 of the absolute source path). All legacy DLC comparisons are valid.
  • Legacy threshold 0.5 sits far below the DLC operating point: cross-fold EER thresholds are 0.87-0.998; median bona fide P(attack) is 0.28-0.96. Scores shift up on an unseen domain.
  • BPCER at the calibrated threshold is concentrated by document type (L6: alb_id 0.32, grc_passport 0.25, fin_id 0.17 ... est_id 0.00); APCER worst on esp/fin/est/svk ids (~0.19-0.21).
  • Gate arithmetic: ACER <= 0.08 needs EER <= ~0.08; across these models that is AUC >= ~0.97.
  • L1 (block 1_2) scores 0.786 here vs 0.717 historically: it was validated with print attacks included; here screen-only. L9 ConvNeXt 0.845 vs 0.854 (pipeline detail: on-the-fly pad).

exp_003 diagnosis - dataset shortcuts (2026-09-24)

Measured on 800/800 even draws per dataset (source crop size, before fitting):

dataset portrait bona / attack median long side bona / attack
DLC 0.03 / 0.00 1298 / 1866 px
SaudiDocs 0.19 / 0.48 826 / 1080 px
KID34K 0.35 / 0.36 892 / 915 px
  • SaudiDocs carries an orientation cue, DLC a resampling cue (attacks shrunk ~2.1x onto the 896 canvas, bona ~1.45x -> different sharpness). KID34K is clean on both, which is why a model trained on DLC+Saudi collapses on it (0.69) while KID+Saudi -> DLC looks fine (0.92; possibly flattered by the same DLC cue).
  • exp_003 on KID: portrait AUC 0.640 vs landscape 0.740, and a large score offset between them (median bona P(attack) 0.21 portrait vs 0.73 landscape) that wrecks pooled ranking.
  • Per-subject medians show ranking failures inside single users (e.g. U10 bona 0.36 / attack 0.30, U13 0.72 / 0.03), so orientation is not the whole story.
  • Fixes queued: ORIENT=landscape (exp_014/015). The resampling cue has no clean fix under the never-upscale rule yet; candidates: random pre-downscale of every training image so sharpness stops tracking the label, or a native-resolution patch input.

Visual check (2026-09-24, contact sheet of fitted images)

  • DLC screen attacks are blatant: over-saturated colours, glare, one shows a mouse cursor. KID34K screen attacks are tight card crops with subtle cues. A model trained on DLC learns "vivid" and fails on KID.
  • SaudiDocs "cropped" images keep ~1.4x margin around the card; backgrounds differ by class (bona: rocks, hands, fabric, wood; screen: keyboards, certificates, Arabic text pages - the background of the displayed photo). Same single card throughout, so background is an easy shortcut. crop_report.csv boxes locate the card in most images (some misses: hand shots, a passport page) -> SAUDI_TIGHT knob. Report lists some files as .JPG where the disk has .jpg: match case-insensitively.

KID-fold ceiling (2026-09-24 04:00)

Five different single variables on the KID fold (baseline, head init, orientation, tight Saudi crop) all land at mean-epoch AUC 0.664-0.676, EER 0.35-0.37; every run reaches train AUC 1.000 in epoch 1. The legacy 46-trial sweep (DLC -> KID, no Saudi) also topped at 0.694. Working hypothesis: DLC + a single Saudi card do not contain the kind of screen recapture KID34K has (subtle, tight phone crops vs DLC's saturated/glare recaptures), i.e. a data ceiling rather than a model problem. Remaining model-side tests: frozen foundation features (exp_021), less fine-tuning (exp_020), patches (exp_016), ConvNeXt (exp_017).

Flag for the owner (not acted on): if these also stay near 0.69, the "worst held-out fold >= 0.90 AUC" gate is not reachable with DLC + SaudiDocs as the only sources for the KID fold. Options to discuss: add a third training domain (e.g. BID/RID) so every fold trains on

= 2 public domains; or report the KID fold as a stress test and gate on the DLC fold + SaudiDocs smoke fold. The gate is unchanged until the owner decides.

BID/RID checked as a third training domain (2026-09-24 06:30) - rejected, not trained

/workspace/cache/kagglehub/datasets/yousef10pc/bid-rid/versions/1/BID RID/: BID = 28,800 bona fide *_in.jpg over 8 doc types (plus _gt_ocr.txt, _gt_segmentation.jpg); RID screen = 216 images (v1-screen-recaptures/screen-recaptures/{iPhone12,iPhone8}), CNH only. Contact sheet: BID bona fide are flat scans (open booklets, any rotation/size); RID screen are camera photos at a fixed 1344x848 of the upper half of the document. Class is confounded with capture pipeline and image size, and 216 attacks add almost no screen diversity. Using it would teach "scan vs photo". Do not use for screen PAD training.

Owner decision + colour cue (2026-09-24 07:00)

  • Owner: if KID-held-out stays poor after reasonable attempts, gate on DLC + SaudiDocs folds only and keep KID as a stress test. Free hand on methods.
  • Post-hoc ensembles on the identical DLC val set (exp_001 legacy scores + exp_002 + exp_006): best pair exp_002+L6 0.950 / EER 0.114, all 9 -> 0.940. No gain: errors are shared.
  • Worst errors of exp_002: DLC bona fide with a strong magenta/purple colour cast scored 0.99 attack (fin_id, svk_id, grc_passport, esp_id, srb_passport); pale screen attacks scored < 0.02. Median bona P(attack) by type: est 0.05, aze 0.08, srb 0.17, esp 0.22, lva 0.33, rus 0.35, svk 0.45, grc 0.46, alb 0.58, fin 0.77. The learned cue is colour cast / saturation.

Parallel-session incident (2026-09-24 11:40)

Two Claude sessions edited queue.txt / ledger / CLAUDE.md at 11:17 and created duplicate numbers for the same ConvNeXt follow-ups. Resolved: kept exp_027 (Saudi fold), exp_028 (EMA), exp_029 (dinov3 DLC), exp_030 (convnext_small DLC), exp_031 (seed 7); the parallel session's colour-aug idea renumbered to exp_032_convnext_wbaug_valDLC; never-run duplicates exp_028_convnext_valDLC_seed7 and exp_029_convnext_ema_valDLC deleted locally (their notebook.ipynb may exist on HF - ignore those two folders).

New gates applied (2026-09-24 16:00)

  • DLC fold, exp_008: AUC 0.973, cf ACER 0.082, BPCER 0.081, acc 0.918 -> PASS.
  • SaudiDocs fold, exp_027: AUC 0.933 (pass). Leave-one-subject-out EER threshold: yousef->fahad ACER 0.145 (APCER 0.258, BPCER 0.033, acc 0.774); fahad->yousef ACER 0.163 (APCER 0.100, BPCER 0.225, acc 0.850). Oracle min ACER on all Saudi 0.143 -> FAIL (ACER). Binding problem: fahad's screen captures (median P(attack) 0.68 vs yousef 0.97).

Gates corrected by owner (2026-09-24 17:30)

AUC >= 0.88, APCER <= 0.10, BPCER <= 0.10, accuracy >= 0.85, soft EER <= 0.15 (ideally 0.10); no ACER gate. DLC exp_028: APCER 0.078 BPCER 0.074 acc 0.924 EER 0.075 -> PASS (EER <= 0.10). Saudi exp_027: EER 0.149; leave-one-subject-out APCER/BPCER 0.258/0.033 and 0.100/0.225 -> FAIL.

Threshold procedure (owner, 2026-09-24 17:45)

Per-epoch validation at the fixed 0.5; after training, an extra validation pass picks the data threshold (EER point of the best checkpoint) - reported as at_val_threshold and saved to threshold.json. At that threshold: DLC exp_028 thr 0.9705 -> APCER 0.075, BPCER 0.075, acc 0.925 (PASS); Saudi exp_027 thr 0.137 -> APCER 0.149, BPCER 0.149, acc 0.851 (FAIL).

Two-tier gates (owner, 2026-09-24 18:30)

  1. Generalisation gate (held-out datasets DLC, SaudiDocs; KID = stress test): AUC >= 0.88, APCER <= 0.10, BPCER <= 0.10, acc >= 0.85, EER <= 0.15 (ideally <= 0.10).
  2. Production gate (new people/devices from a domain that IS in training): APCER <= 1%, BPCER <= 1%, EER <= 1.5%, AUC >= 0.98, accuracy >= 98%. Final measurement needs a new Saudi capture set of >= 20-30 people not used in training (does not exist yet). Proxy until then: exp_034/035 (owner-approved exception: KID / DLC split by subject, whole subjects per side). Why 1% is not a held-out-dataset gate: best cross-domain EER 0.075 (DLC) / 0.149 (Saudi); and the DLC val set has 80 documents x ~62 frames, so one bad document alone is ~1.2% BPCER.

SaudiDocs error analysis (2026-09-25 00:00, exp_027 scores)

  • False rejects: genuine cards lying on laptop keyboards, next to a laptop screen, on a purple folder -> P(attack) 0.97-0.98. The model uses the card's surroundings (keyboards/screens). -> exp_038 crops the held-out Saudi images to the card (SAUDI_TIGHT).
  • Missed attacks cluster by recording: 10 of 100 screen videos give 37% of all misses (Screen_yousef_54/88/87/85 ...: 60-80% of their frames missed); 4 of 64 genuine videos have

    50% of frames rejected. The Saudi fold is effectively ~164 videos, not 5,591 images - its error bars are wide, and new Saudi data should add sessions/devices, not frames per video.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support