Overview
DEIMv2 is an evolution of the DEIM (DETR with Improved Matching) framework, extended with rich features from DINOv3. DEIM's core contribution β Dense One-to-One (Dense O2O) label assignment β accelerates convergence of DETR-style detectors versus the traditional sparse one-to-one matching used in DETR/Deformable-DETR, without sacrificing the end-to-end, NMS-free detection pipeline.
DEIMv2 spans eight model sizes from ultra-light (Atto) to extra-large (X), covering GPU, edge, and mobile deployment budgets. For the X/L/M/S variants, DEIMv2 adopts DINOv3-pretrained or DINOv3-distilled ViT backbones and introduces a Spatial Tuning Adapter (STA) that converts DINOv3's single-scale output into multi-scale features, complementing strong semantics with fine-grained spatial detail. The ultra-lightweight variants (N/Pico/Femto/Atto) instead use a depth- and width-pruned HGNetv2 backbone to meet strict resource budgets. Combined with a simplified decoder and an upgraded Dense O2O scheme, DEIMv2 achieves a strong performance-cost trade-off across the board, with the deimv2_s model notably surpassing 50 AP on the challenging COCO benchmark at under 10M parameters.
License note: DEIMv2 is released by Intellindust AI Lab under a non-commercial research license (see LICENSE.md) β commercial use requires a separate license from Intellindust. Review the upstream license terms before deploying these weights in a commercial product.
Model Variants
| Model | Backbone | Input Size | Params(M) | mAP[.5:.95]% | Validated Devices | Config |
|---|---|---|---|---|---|---|
deimv2_atto |
HGNetv2-Atto | 320Γ320 | 0.5 | 23.8 | N/A | N/A |
deimv2_femto |
HGNetv2-Femto | 416Γ416 | 1.0 | 31.0 | N/A | N/A |
deimv2_pico |
HGNetv2-Pico | 640Γ640 | 1.5 | 38.5 | N/A | N/A |
deimv2_n |
HGNetv2-B0 | 640Γ640 | 3.6 | 43.0 | N/A | N/A |
deimv2_s |
DINOv3-vit_tiny | 640Γ640 | 9.7 | 50.9 | TDA4VH | deimv2_s_config.yaml |
deimv2_m |
DINOv3-vit_tinyplus | 640Γ640 | 18.1 | 53.0 | TDA4VH | deimv2_m_config.yaml |
deimv2_l |
DINOv3-vit_small | 640Γ640 | 32.2 | 56.0 | N/A | N/A |
deimv2_x |
DINOv3-vit_small+ | 640Γ640 | 50.3 | 57.8 | N/A | N/A |
mAP values are on COCO val2017, as reported by the upstream DEIMv2 repository.
Recommended for edge deployment: deimv2_s (best accuracy/compute trade-off; the only variant marked recommended: true in its TIDL config)
Quick Start
Prerequisites
# Install core dependencies (auto-installed by prepare_model.py if missing)
pip install torch>=1.12.0 torchvision>=0.13.0 onnx>=1.14.0 huggingface_hub timm calflops
# ONNX inference
pip install onnxruntime>=1.15.0
# scipy is required because DEIM imports it at module load time (models/matcher.py)
pip install scipy
Note:
prepare_model.pyuseshuggingface_hub, which requires git and internet access on first use to clone the DEIMv2 source and download pretrained weights (10β200 MB from HuggingFace Hub). Subsequent runs reuse the cache at `/.cache/deimv2_srcand~/.cache/huggingface/hub`.
Export the Model
Pretrained COCO weights are downloaded automatically via huggingface_hub on first use.
# List all available variants with accuracy info
python prepare_model.py --list-models
# Export the default model (deimv2_s)
python prepare_model.py
# Export a specific model variant
python prepare_model.py --model deimv2_m
# Export multiple variants at once
python prepare_model.py --model deimv2_s deimv2_m deimv2_l
# Export all supported models
python prepare_model.py --model all
# Export with a custom input resolution
python prepare_model.py --model deimv2_s --shape 800 800
# Export from a locally trained checkpoint
python prepare_model.py --model deimv2_s --weights /path/to/checkpoint.pth
# Force re-export even if the .onnx already exists
python prepare_model.py --model deimv2_s --force
The script automatically:
- Installs missing dependencies (
torch,onnx,huggingface_hub,timm,scipy) if not present - Clones the DEIMv2 source repository via git on first use (cached at
~/.cache/deimv2_src) - Downloads the pretrained COCO weights for the requested variant(s) from HuggingFace Hub
- Wraps the model (backbone + encoder + decoder, skipping the postprocessor) to accept a plain
(N, 3, H, W)tensor - Exports to ONNX (opset 17 by default) with constant folding and shape inference
- Simplifies the graph with
onnxslim/onnxsim(best-effort) and saves<model_key>.onnxin the output directory
ONNX model inputs / outputs:
| Tensor | Shape | Description |
|---|---|---|
images (input) |
(N, 3, H, W) |
ImageNet-normalized float32 |
pred_boxes (output 0) |
(N, num_queries, 4) |
Boxes as (cx, cy, w, h), normalized [0, 1] |
pred_logits (output 1) |
(N, num_queries, 80) |
Raw class logits for 80 COCO classes |
DEIMv2 outputs a fixed number of query slots per image (100β300 depending on variant) regardless of the number of objects present.
Input preprocessing β DEIMv2 expects ImageNet-normalized inputs:
import cv2
import numpy as np
mean = np.array([123.675, 116.28, 103.53], dtype=np.float32)
scale = np.array([0.017125, 0.017507, 0.017429], dtype=np.float32) # 1/255 / std
img = cv2.imread("image.jpg") # BGR uint8
h, w = MODEL_SHAPE # e.g. (640, 640) for S/M/L/X; (320,320) atto; (416,416) femto
img = cv2.resize(img, (w, h))
img = img.astype(np.float32)
img = (img - mean) * scale
img = np.transpose(img, (2, 0, 1)) # HWC β CHW
img = np.expand_dims(img, 0) # add batch dim β (1, 3, H, W)
Post-processing β class scores are computed via sigmoid (not softmax):
import numpy as np
CONFIDENCE_THRESHOLD = 0.25
def postprocess(pred_boxes, pred_logits, image_h, image_w, threshold=CONFIDENCE_THRESHOLD):
boxes = pred_boxes[0] # (num_queries, 4) cx,cy,w,h normalized
logits = pred_logits[0] # (num_queries, 80)
scores = 1 / (1 + np.exp(-logits)) # sigmoid
scores = scores.max(axis=1)
labels = scores.argmax(axis=1)
keep = scores > threshold
cx, cy, bw, bh = boxes[keep].T
x1 = (cx - bw / 2) * image_w
y1 = (cy - bh / 2) * image_h
x2 = (cx + bw / 2) * image_w
y2 = (cy + bh / 2) * image_h
return np.stack([x1, y1, x2, y2], axis=1), labels[keep], scores[keep]
Compile and Infer uing edgeai-tidlrunner
Note: Run the commands below from inside the
tidlrunnerdirectory (the cloned edgeai-tidlrunner repository), with--config_pathpointing to this model's config file.
Compile using edgeai-tidlrunner - on PC
cd /path/to/edgeai-tidlrunner
tidlrunner-cli compile --target_device J784S4 \
--config_path /path/to/deimv2_s_config.yaml
Run Inference Benchmark - on device
cd /path/to/edgeai-tidlrunner
tidlrunner-cli infer --target_device J784S4 \
--config_path /path/to/deimv2_s_config.yaml
Replace
deimv2_s_config.yamlwithdeimv2_m_config.yamlto compile/infer thedeimv2_mvariant. To evaluate accuracy instead of just compiling, replacecompilewithevaluate.
Compile and Infer using edgeai-tidl-tools (Advanced):
Follow the instructions at https://github.com/TexasInstruments/edgeai-tidl-tools
Deploy using edgeai-tidl-tools:
Deplyment can be done using edgeai-tidl-tools. For ONNX models, onnxruntime-tidl with TIDL acceleration can be used. Consult the documentation of edgeai-tidl-tools for more details.
Citation
If you use these models, please cite:
@article{huang2025deimv2,
title = {Real-Time Object Detection Meets DINOv3},
author = {Huang, Shihua and Hou, Yongjie and Liu, Longfei and Yu, Xuanlong and Shen, Xi},
journal = {arXiv preprint arXiv:2509.20787},
year = {2025}
}
π Resources
| Resource | Link |
|---|---|
| Paper | arXiv:2509.20787 |
| Source Code | Intellindust-AI-Lab/DEIMv2 |
| License | LICENSE.md (non-commercial) |
| HGNetv2 Backbone | Peterande/HGNetv2 |
| DINOv3 Backbone | facebookresearch/dinov3 |
| COCO Dataset | cocodataset.org |
| edgeai-tidl-tools | GitHub |
| edgeai-tidlrunner | GitHub |
| EdgeAI SDK | Documentation |
Related Models
|
DETR Original end-to-end DETR transformer detector |
Deformable-DETR Deformable attention for faster convergence |
RT-DETRv2 Real-time DETR transformer detector |
RF-DETR Real-time DETR with flexible backbones |
Maintained by: Texas Instruments EdgeAI Team
Last Updated: August 2026