ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection
Tongtong Wang1 Mingzhu Xu1β Chenglong Yu1 Jing Wang1 Xiaohui Lin1 Weili Guan2
1School of Software, Shandong University
2Harbin Institute of Technology, Shenzhen
βCorresponding author
π Model Description
This repository provides the official model checkpoints for ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection, accepted by ACM Multimedia 2026.
Infrared Small Target Detection aims to accurately segment weak and tiny targets from complex infrared backgrounds. Existing pure-vision methods rely mainly on pixel-level information, while existing vision-language methods commonly describe targets and backgrounds using a single textual prompt. Such a symmetric design overlooks the inherent semantic differences between sparse infrared targets and structurally complex backgrounds.
ADGNet addresses this problem through three main components:
- Asymmetric Dual-text Prompt (ADP): uses an abstract, image-independent target prompt and a detailed, image-dependent background prompt.
- Asymmetric Dual-Branch Interaction (ADBI): independently performs target localization and background suppression using their corresponding textual priors.
- Adaptive Feature Aggregation (AFA): dynamically fuses target-enhanced and background-suppressed features for accurate segmentation.
The model uses the pretrained CLIP ViT-B/16 text encoder to extract semantic representations from the target and background prompts.
π Available Checkpoints
All ADGNet checkpoints are hosted in this Hugging Face model repository.
Download the required checkpoint directly from the Files and versions section of this repository.
| Dataset | Checkpoint |
|---|---|
| IRSTD-1K | ADGNet_mIoU_72.38_IRSTD-1K.pth.tar |
| NUDT-SIRST | ADGNet_mIoU_95.53_NUDT-SIRST.pth.tar |
| SIRST | ADGNet_mIoU_83.08_SIRST.pth.tar |
π Usage
These checkpoints are designed to be used with the official ADGNet implementation:
https://github.com/iLearn-Lab/MM26-ADGNet
1. Clone the Official Repository
git clone https://github.com/iLearn-Lab/MM26-ADGNet.git
cd MM26-ADGNet
2. Prepare the Checkpoints
Place the downloaded checkpoints in:
MM26-ADGNet/
βββ SOTA_pth/
βββ ADGNet_mIoU_72.38_IRSTD-1K.pth.tar
βββ ADGNet_mIoU_95.53_NUDT-SIRST.pth.tar
βββ ADGNet_mIoU_83.08_SIRST.pth.tar
3. Prepare the CLIP Text Encoder
ADGNet uses the pretrained CLIP ViT-B/16 model:
git clone https://huggingface.co/openai/clip-vit-base-patch16
Update the local CLIP model path in the corresponding project configuration or source file before inference.
4. Run Evaluation
Example evaluation on IRSTD-1K:
python train.py \
--trainset "IRSTD-1K" \
--testset "IRSTD-1K" \
--dataset_dir "./datasets" \
--mode test \
--ckpt "./SOTA_pth/ADGNet_mIoU_72.38_IRSTD-1K.pth.tar"
Replace the dataset name and checkpoint path when evaluating on NUDT-SIRST or SIRST.
π Dataset and Text Annotation Preparation
The original infrared images and ground-truth masks are not included in this model repository. Please obtain IRSTD-1K, NUDT-SIRST, and SIRST from their respective official sources.
The asymmetric text annotations used by ADGNet are released separately in our Hugging Face dataset repository:
- AITIR Text Annotations:
Download
After downloading the original datasets and text annotations, organize them according to the official ADGNet repository:
datasets/
βββ IRSTD-1K/
β βββ images/
β βββ masks/
β βββ img_idx/
β βββ text/
βββ NUDT-SIRST/
β βββ images/
β βββ masks/
β βββ img_idx/
β βββ text/
βββ SIRST/
βββ images/
βββ masks/
βββ img_idx/
βββ text/
π― Intended Use
The released checkpoints are intended for:
- Academic research on infrared small target detection
- Reproduction of the results reported in the ADGNet paper
- Evaluation on IRSTD-1K, NUDT-SIRST, and SIRST
- Research on multimodal and text-guided infrared image segmentation
- Comparison with other infrared small target detection methods
β οΈ Limitations
- The model requires both infrared images and corresponding textual prompts.
- Detection performance may vary when applied to datasets or scenes that differ substantially from the training distribution.
- The released checkpoints are designed for the dataset splits and evaluation settings used in the paper.
- The model depends on the pretrained CLIP ViT-B/16 text encoder.
- The original infrared datasets are subject to their respective licenses and terms of use.
π Related Resources
- GitHub Repository: iLearn-Lab/MM26-ADGNet
- Paper:
ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection - Text Annotations:
AITIR Text Annotations
π Citation
If you find ADGNet or the released checkpoints useful in your research, please consider citing our paper:
Please also consider checking out and citing our other related work: