EfficientNet-B0 Video Deepfake Detector

Lightweight video-level real/fake classifier using ImageNet-pretrained EfficientNet-B0.

Model

  • Architecture: EfficientNet-B0
  • Classes: REAL, FAKE
  • Frames per video: 16
  • Frame size: 224x224
  • Temporal aggregation: mean pooling of frame logits
  • Input: video file

Preprocessing

  1. Uniformly sample 16 frames.
  2. Resize frames to 224x224.
  3. Apply ImageNet normalization.
  4. Run EfficientNet-B0 on every frame.
  5. Mean-pool frame logits for the video prediction.

See preprocessing.py for the exact preprocessing implementation.

Repository files

  • pytorch_model.bin - trained weights
  • config.json - model and preprocessing metadata
  • preprocessing.py - video preprocessing
  • requirements.txt - runtime dependencies
  • README.md - model card

Loading

import json
import torch
import torch.nn as nn
from torchvision.models import efficientnet_b0

with open("config.json") as f:
    config = json.load(f)

model = efficientnet_b0(weights=None)
model.classifier[1] = nn.Linear(
    model.classifier[1].in_features,
    config["num_classes"]
)
model.load_state_dict(
    torch.load("pytorch_model.bin", map_location="cpu")
)
model.eval()

This is a PyTorch/torchvision model repository, not a native Transformers model.

Downloads last month
18
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using iron04/deepfake_efficientnet 1