Understand Before Detect: Vision--Language Learning for Omni-Domain Infrared Small Target Detection

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the performance limitations of small target detection in heterogeneous infrared imaging domains, which stem from domain shift and insufficient semantic information. To overcome these challenges, the paper introduces a novel "understanding-before-detection" paradigm by integrating vision–language learning into universal infrared small target detection for the first time. The approach leverages language supervision to establish global target understanding, incorporates a language-aligned cross-domain representation mechanism, and introduces a low-rank semantic interaction module. Additionally, the authors construct OmniIRST-VL, the first large-scale multimodal infrared dataset. Through joint vision–language learning and multi-task instruction fine-tuning, the proposed method achieves strong cross-domain generalization within a single model and significantly outperforms existing approaches on OmniIRST-VL.
📝 Abstract
Omni-domain infrared small target (IRST) detection is crucial for infrared surveillance, yet remains challenging due to heterogeneous imaging domains and inconsistent target characteristics. Previous deep learning-based methods have been developed for visual-only paradigms and achieved promising performance on domain-specific tasks. However, existing methods follow the task-specific supervised learning paradigm. This paradigm simplifies the full-scene infrared observations to sparse target supervision, discarding the semantics that remain invariant across heterogeneous domains. Consequently, detection performance suffers substantially under domain shifts. To handle this issue, we introduce \textbf{``understand before detect''}, a paradigm that formulates omni-domain IRST detection as an understanding-driven process, where holistic infrared target understanding precedes precise detection. Building on this paradigm, we propose \textbf{JinSight}, which first develops holistic IRST understanding through language supervision and then transfers the learned cross-domain representations to precise small-target detection. By grounding infrared representations in language semantics, JinSight enables a single model to generalize across heterogeneous infrared domains. We then introduce Latent Semantic Interaction (LSI), which exchanges language-aligned global semantics with fine-grained spatial features in a compact low-rank space. To address the lack of multimodal omni-domain IRST benchmarks, we build \textbf{OmniIRST-VL}, the first large-scale, highly diverse vision--language dataset for omni-domain IRST detection. It comprises over 39k annotations across six complementary instruction tasks covering both scene-level understanding and target-centric reasoning.
Problem

Research questions and friction points this paper is trying to address.

infrared small target detection
omni-domain
domain shift
vision-language learning
semantic understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language learning
omni-domain detection
infrared small target
semantic grounding
latent semantic interaction
H
Haoyang Yuan
National University of Defense Technology, Changsha, China
Boyang Li
Boyang Li
National University of Defense Technology
Infrared small target detectionWeakly supervised semantic segmentation
Yingqian Wang
Yingqian Wang
National University of Defense Technology
light fieldimage super-resolution
Y
Yimian Dai
College of Computer Science, Nankai University, Tianjin, China
N
Nuo Chen
Peking University, Beijing, China
X
Xinfei Huang
National University of Defense Technology, Changsha, China
S
Shuqi Yi
National University of Defense Technology, Changsha, China
Z
Zaiping Lin
National University of Defense Technology, Changsha, China
W
Weidong Sheng
National University of Defense Technology, Changsha, China
W
Wei An
National University of Defense Technology, Changsha, China