MADL model weights

This repository is the weight companion for MADL: Towards Dependable Image Forgery Detection via Multi-Agent Forensic Reasoning. The executable code, configuration files, and evaluation scripts are distributed in the companion GitHub code repository.

Model repository: moindy/MADL.

MADL produces one of three image-level labelsβ€”real, synthetic, or tamperedβ€”and returns a binary localization mask only for locally tampered images. The release contains the trainable components developed for MADL; it does not redistribute the Qwen base model or SAM ViT-H.

Files

Directory Component Purpose
agent_a_qwen_lora/ Agent A LoRA adapter Adapts Qwen2.5-VL-7B-Instruct to three-class semantic verification and region proposals.
agent_b_dualstream/ Agent B dual-stream checkpoint RGB/SRM ConvNeXt-Tiny classifier and continuous manipulation heatmap.
agent_b_visual_ranker/ Agent B visual-ranker checkpoint Selects a final SAM candidate using RGB, mask, heatmap, and geometry evidence.
MANIFEST.json Integrity manifest Records the byte size and SHA-256 of every released model artifact.

The dual-stream release checkpoint contains only the model state, architecture, class names, and a compact training summary. Optimizer state and local training paths were intentionally removed. The visual-ranker checkpoint uses the strict four-field runtime schema required by the companion loader.

External dependencies

  1. Download Qwen/Qwen2.5-VL-7B-Instruct through Transformers or the Hugging Face Hub. PEFT reads this identifier from adapter_config.json.
  2. Download the official SAM ViT-H checkpoint sam_vit_h_4b8939.pth from Meta's Segment Anything release.
  3. Place this repository under the runtime weight root and place SAM at external/sam_vit_h_4b8939.pth, or override the relative paths in the MADL YAML configuration.

Expected layout:

models/
β”œβ”€β”€ agent_a_qwen_lora/
β”œβ”€β”€ agent_b_dualstream/
β”‚   └── madl_agent_b_dualstream_v1.pt
β”œβ”€β”€ agent_b_visual_ranker/
β”‚   └── madl_agent_b_visual_ranker_v1.pt
└── external/
    └── sam_vit_h_4b8939.pth

Verify the files before inference:

python scripts/verify_weights.py /path/to/models

Training data and evaluation scope

The released components were trained for the SID-Set task, which distinguishes authentic, fully synthetic, and locally tampered images and provides masks for the locally tampered class. The release does not include SID-Set images. Users must obtain the dataset under its own terms.

Intended use

The weights are intended for academic research on image forgery detection, localization, evidence interaction, and multi-agent forensic reasoning. They are not a substitute for human forensic examination and should not be used as the sole basis for legal, disciplinary, or content-removal decisions.

Limitations

  • Very small, fragmented, or weak-trace manipulations remain difficult.
  • Cross-dataset behavior has not been established by the final manuscript.
  • Multiple model calls increase latency and GPU memory requirements.
  • The natural-language field is a structured evidence trace; the release does not claim that it has been independently evaluated as explanation quality.
  • Results depend on the exact Qwen, SAM, preprocessing, and threshold versions documented by the companion code.

Licenses and attribution

The released MADL code and original weight packaging are provided under Apache-2.0. Qwen2.5-VL, SAM, SID-Set, and other dependencies retain their own licenses and terms. The code README lists the relevant upstream projects. Dataset-derived use should preserve SID-Set attribution.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for moindy/MADL

Finetuned
(1250)
this model