🤖 AI Summary
This work addresses the performance limitations of small target detection in heterogeneous infrared imaging domains, which stem from domain shift and insufficient semantic information. To overcome these challenges, the paper introduces a novel "understanding-before-detection" paradigm by integrating vision–language learning into universal infrared small target detection for the first time. The approach leverages language supervision to establish global target understanding, incorporates a language-aligned cross-domain representation mechanism, and introduces a low-rank semantic interaction module. Additionally, the authors construct OmniIRST-VL, the first large-scale multimodal infrared dataset. Through joint vision–language learning and multi-task instruction fine-tuning, the proposed method achieves strong cross-domain generalization within a single model and significantly outperforms existing approaches on OmniIRST-VL.
📝 Abstract
Omni-domain infrared small target (IRST) detection is crucial for infrared surveillance, yet remains challenging due to heterogeneous imaging domains and inconsistent target characteristics. Previous deep learning-based methods have been developed for visual-only paradigms and achieved promising performance on domain-specific tasks. However, existing methods follow the task-specific supervised learning paradigm. This paradigm simplifies the full-scene infrared observations to sparse target supervision, discarding the semantics that remain invariant across heterogeneous domains. Consequently, detection performance suffers substantially under domain shifts. To handle this issue, we introduce \textbf{``understand before detect''}, a paradigm that formulates omni-domain IRST detection as an understanding-driven process, where holistic infrared target understanding precedes precise detection. Building on this paradigm, we propose \textbf{JinSight}, which first develops holistic IRST understanding through language supervision and then transfers the learned cross-domain representations to precise small-target detection. By grounding infrared representations in language semantics, JinSight enables a single model to generalize across heterogeneous infrared domains. We then introduce Latent Semantic Interaction (LSI), which exchanges language-aligned global semantics with fine-grained spatial features in a compact low-rank space. To address the lack of multimodal omni-domain IRST benchmarks, we build \textbf{OmniIRST-VL}, the first large-scale, highly diverse vision--language dataset for omni-domain IRST detection. It comprises over 39k annotations across six complementary instruction tasks covering both scene-level understanding and target-centric reasoning.