EfficientNet-B3 Image Forgery Detection

This repository contains a fine-tuned EfficientNet-B3 model for binary classification of synthetic identity-document images as genuine or forged.

The checkpoint is provided as a PyTorch state_dict:

efficientnet_b3_finetuned.pt

Model description

  • Architecture: EfficientNet-B3
  • Framework: PyTorch and timm
  • Task: binary image classification
  • Input: one identity-document image
  • Output: genuine or forged probability
  • Input resolution: 300 × 300 pixels

Labels:

0 = genuine / bona fide
1 = forged / tampered

Dataset

The model was trained and evaluated on the template/composite portion of the Synthetic dataset of ID and Travel Documents (SIDTD), generated using MIDV-2020 document templates.

Dataset composition used in this project:

Genuine images: 1,000
Forged images:  1,222
Total images:   2,222

The forged samples include the following SIDTD annotation types:

  • Crop_and_Replace
  • Inpaint_and_Rewrite

The dataset itself is not included in this repository. Please follow the SIDTD dataset terms and citation requirements before downloading or redistributing it.

Preprocessing

Each image is:

  1. converted to RGB;
  2. resized to 300 × 300 pixels;
  3. converted to a tensor;
  4. normalized using ImageNet mean and standard deviation:
mean = [0.485, 0.456, 0.406]
std  = [0.229, 0.224, 0.225]

Training used mild brightness, contrast, and saturation variation. Horizontal flipping was not used because mirrored identity documents are not realistic examples.

Training

The model was initialized from a pretrained EfficientNet-B3. Phase 1 trained the binary classifier head. The final checkpoint was then fine-tuned by unfreezing the final EfficientNet feature block and classifier head.

Fine-tuning used:

  • optimizer: AdamW;
  • learning rate: 1e-4;
  • weight decay: 1e-4;
  • epochs: 20;
  • binary cross-entropy equivalent: two-class cross-entropy loss;
  • deterministic stratified hold-out split: 80% train, 10% validation, 10% test.

Evaluation

On the fixed 223-image test split at a decision threshold of 0.5:

Accuracy:  98.21%
Precision: 97.60%
Recall:    99.19%
F1:        98.39%
ROC-AUC:   99.98%

Confusion matrix, with rows representing actual labels and columns representing predictions:

[[97,   3],
 [ 1, 122]]

Subtype recall on the forged test samples:

Crop-and-replace: 94.74%
Inpainting:       100.00%

These results are specific to the project’s synthetic dataset and split. They should not be interpreted as real-world identity-verification performance.

Usage

Install the required packages:

pip install torch torchvision timm pillow

Load the checkpoint:

from pathlib import Path

import timm
import torch
from PIL import Image
from torchvision import transforms


device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

model = timm.create_model(
    "efficientnet_b3",
    pretrained=False,
    num_classes=2,
)
model.load_state_dict(
    torch.load(
        "efficientnet_b3_finetuned.pt",
        map_location="cpu",
        weights_only=True,
    )
)
model.to(device)
model.eval()

preprocess = transforms.Compose([
    transforms.Resize((300, 300)),
    transforms.ToTensor(),
    transforms.Normalize(
        mean=(0.485, 0.456, 0.406),
        std=(0.229, 0.224, 0.225),
    ),
])

image = Image.open("document.jpg").convert("RGB")
input_tensor = preprocess(image).unsqueeze(0).to(device)

with torch.no_grad():
    probabilities = torch.softmax(model(input_tensor), dim=1)[0]

genuine_probability = float(probabilities[0])
forged_probability = float(probabilities[1])
label = "forged" if forged_probability >= 0.5 else "genuine"

print({
    "label": label,
    "genuine_probability": genuine_probability,
    "forged_probability": forged_probability,
})

Explainability

The project uses Grad-CAM with the final spatial convolutional layer of EfficientNet-B3. The resulting heatmap indicates image regions that contributed to the selected prediction.

Grad-CAM is an explanation aid and should not be interpreted as a guaranteed localization of the tampered region.

Intended use

This model is intended for:

  • educational computer-vision projects;
  • research prototypes;
  • synthetic identity-document forgery experiments;
  • demonstrations of explainable image classification;
  • offline error analysis.

Limitations

  • The training data is synthetic.
  • The model has not been validated on real identity documents.
  • Performance may depend on document templates, image quality, and dataset artifacts.
  • The binary output is not equivalent to full identity verification.
  • A positive or negative result should not be used as the sole basis for a KYC or access-control decision.
  • The reported test split is an image-level random split and may not measure generalization to unseen document families as strictly as a grouped or external split.

Citation and attribution

This model uses the SIDTD dataset and MIDV-2020-derived document templates. Please cite and follow the terms of the original SIDTD and MIDV-2020 resources when using this model or reproducing the experiments.

SIDTD project: https://github.com/Oriolrt/SIDTD_Dataset

License

The model repository is released under the MIT license. Dataset licensing, attribution, and redistribution terms remain subject to the original SIDTD and MIDV-2020 sources.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support