---
license: apache-2.0
pipeline_tag: image-segmentation
tags:
- infrared-small-target-detection
- multimodal
- vision-language
- pytorch
---
ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection
Tongtong Wang1
Mingzhu Xu1✉
Chenglong Yu1
Jing Wang1
Xiaohui Lin1
Weili Guan2
1School of Software, Shandong University
2Harbin Institute of Technology, Shenzhen
✉Corresponding author
## 📖 Model Description
This repository provides the official model checkpoints for **ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection**, accepted by **ACM Multimedia 2026**.
Infrared Small Target Detection aims to accurately segment weak and tiny targets from complex infrared backgrounds. Existing pure-vision methods rely mainly on pixel-level information, while existing vision-language methods commonly describe targets and backgrounds using a single textual prompt. Such a symmetric design overlooks the inherent semantic differences between sparse infrared targets and structurally complex backgrounds.
ADGNet addresses this problem through three main components:
- **Asymmetric Dual-text Prompt (ADP):** uses an abstract, image-independent target prompt and a detailed, image-dependent background prompt.
- **Asymmetric Dual-Branch Interaction (ADBI):** independently performs target localization and background suppression using their corresponding textual priors.
- **Adaptive Feature Aggregation (AFA):** dynamically fuses target-enhanced and background-suppressed features for accurate segmentation.
The model uses the pretrained **CLIP ViT-B/16** text encoder to extract semantic representations from the target and background prompts.
---
## 🏆 Available Checkpoints
All ADGNet checkpoints are hosted in this Hugging Face model repository.
Download the required checkpoint directly from the **Files and versions** section of this repository.
| Dataset | Checkpoint |
| :--------: | :----------------------------------------------------------: |
| IRSTD-1K | [`ADGNet_mIoU_72.38_IRSTD-1K.pth.tar`](https://huggingface.co/iLearn-Lab/MM26-ADGNet/blob/main/ADGNet_mIoU_72.38_IRSTD-1K.pth.tar) |
| NUDT-SIRST | [`ADGNet_mIoU_95.53_NUDT-SIRST.pth.tar`](https://huggingface.co/iLearn-Lab/MM26-ADGNet/blob/main/ADGNet_mIoU_95.53_NUDT-SIRST.pth.tar) |
| SIRST | [`ADGNet_mIoU_83.08_SIRST.pth.tar`](https://huggingface.co/iLearn-Lab/MM26-ADGNet/blob/main/ADGNet_mIoU_83.08_SIRST.pth.tar) |
---
## 🚀 Usage
These checkpoints are designed to be used with the official ADGNet implementation:
```text
https://github.com/iLearn-Lab/MM26-ADGNet
```
### 1. Clone the Official Repository
```bash
git clone https://github.com/iLearn-Lab/MM26-ADGNet.git
cd MM26-ADGNet
```
### 2. Prepare the Checkpoints
Place the downloaded checkpoints in:
```text
MM26-ADGNet/
└── SOTA_pth/
├── ADGNet_mIoU_72.38_IRSTD-1K.pth.tar
├── ADGNet_mIoU_95.53_NUDT-SIRST.pth.tar
└── ADGNet_mIoU_83.08_SIRST.pth.tar
```
### 3. Prepare the CLIP Text Encoder
ADGNet uses the pretrained **CLIP ViT-B/16** model:
```bash
git clone https://huggingface.co/openai/clip-vit-base-patch16
```
Update the local CLIP model path in the corresponding project configuration or source file before inference.
### 4. Run Evaluation
Example evaluation on IRSTD-1K:
```bash
python train.py \
--trainset "IRSTD-1K" \
--testset "IRSTD-1K" \
--dataset_dir "./datasets" \
--mode test \
--ckpt "./SOTA_pth/ADGNet_mIoU_72.38_IRSTD-1K.pth.tar"
```
Replace the dataset name and checkpoint path when evaluating on NUDT-SIRST or SIRST.
---
## 📂 Dataset and Text Annotation Preparation
The original infrared images and ground-truth masks are not included in this model repository. Please obtain **IRSTD-1K**, **NUDT-SIRST**, and **SIRST** from their respective official sources.
The asymmetric text annotations used by ADGNet are released separately in our Hugging Face dataset repository:
- **AITIR Text Annotations:** [`Download`]()
After downloading the original datasets and text annotations, organize them according to the official ADGNet repository:
```text
datasets/
├── IRSTD-1K/
│ ├── images/
│ ├── masks/
│ ├── img_idx/
│ └── text/
├── NUDT-SIRST/
│ ├── images/
│ ├── masks/
│ ├── img_idx/
│ └── text/
└── SIRST/
├── images/
├── masks/
├── img_idx/
└── text/
```
---
## 🎯 Intended Use
The released checkpoints are intended for:
- Academic research on infrared small target detection
- Reproduction of the results reported in the ADGNet paper
- Evaluation on IRSTD-1K, NUDT-SIRST, and SIRST
- Research on multimodal and text-guided infrared image segmentation
- Comparison with other infrared small target detection methods
---
## ⚠️ Limitations
- The model requires both infrared images and corresponding textual prompts.
- Detection performance may vary when applied to datasets or scenes that differ substantially from the training distribution.
- The released checkpoints are designed for the dataset splits and evaluation settings used in the paper.
- The model depends on the pretrained CLIP ViT-B/16 text encoder.
- The original infrared datasets are subject to their respective licenses and terms of use.
---
## 🔗 Related Resources
- **GitHub Repository:** [iLearn-Lab/MM26-ADGNet](https://github.com/iLearn-Lab/MM26-ADGNet)
- **Paper:** [`ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection`]()
- **Text Annotations:** [`AITIR Text Annotations`]()
---
## 📚 Citation
If you find ADGNet or the released checkpoints useful in your research, please consider citing our paper:
```bibtex
```
Please also consider checking out and citing our other related work:
```bibtex
```
---