DAD public-data experiment models

Three models for the public-data experiment in Detect Anything in Graphic Design: Element-Level Rewards for Autoregressive Detection. These are public-data experiment models, separate from the DAD checkpoint used in the paper's main experiments.

Subdirectory Model Selected checkpoint
sft/ Supervised fine-tuning SFT step 5,000
grpo/ GRPO starting from the shared SFT checkpoint RL step 1,000
elerpo/ EleRPO starting from the shared SFT checkpoint RL step 1,000

Each subdirectory contains a complete merged Qwen3-VL-2B model, tokenizer, image processor, and chat template. Select a subdirectory when loading; the repository root contains the shared model card and licenses.

Use

The DAD code repository provides data download, inference, and evaluation:

pip install -r requirements.txt
python -m inference.predict --variant elerpo --output outputs/elerpo.jsonl
python -m inference.evaluate \
  --annotations public_data/annotations/test.jsonl \
  --predictions outputs/elerpo.jsonl

For direct loading with Transformers 5.7 or newer:

from huggingface_hub import snapshot_download
from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
from pathlib import Path
import torch

variant = "elerpo"  # also "sft" or "grpo"
root = snapshot_download("dad887/DAD-Public-Models", allow_patterns=[f"{variant}/*"])
path = Path(root) / variant
processor = AutoProcessor.from_pretrained(path)
model = Qwen3VLForConditionalGeneration.from_pretrained(
    path, dtype=torch.bfloat16, device_map="auto"
)

The models emit semicolon-separated x1,y1,x2,y2,type records, ordered back-to-front. Coordinates are normalized to [0,1000]; t means text and v means visual. The inference code supplies the corresponding image preparation and prompts.

Data and scope

The public experiment dataset supplies fixed SFT, RL, and test splits. Its SFT split draws from released DAD, ChartGalaxy, Crello, PrismLayersPro, and LICA examples. See the paper for training parameters and the dataset card for sample counts and source terms.

These models target graphic-design element detection. They can miss, merge, or split elements and are not general-purpose ground-truth annotators. Sampling, inference backends, batching, and library versions can affect predictions.

License and attribution

The fine-tuned model releases are provided under CC BY-NC 4.0. The original Qwen3-VL-2B-Instruct base model is provided by the Qwen Team under Apache 2.0; its terms and attribution are retained in NOTICE and licenses/Apache-2.0.txt. Dataset source licenses are documented separately in the dataset repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dad887/DAD-Public-Models

Finetuned
(270)
this model

Dataset used to train dad887/DAD-Public-Models