🤖 AI Summary
This study addresses the limited localization accuracy of multimodal large models in infrared small target detection by proposing an agent-based dynamic visual search framework. The method constructs a specialized visual toolchain comprising five complementary tools and designs a zoom-guided interactive learning mechanism. By leveraging annotated trajectory supervision for tool selection and parameter configuration, the framework enables adaptive zooming observation and fine-grained search. Experimental results demonstrate that the proposed approach outperforms existing methods on both the WideIRSTD-Full and IRSTD-1k datasets, significantly enhancing precise localization performance for infrared small targets.
📝 Abstract
Infrared small-target detection plays an important role in maritime monitoring and aerial surveillance. Although multimodal large language models (MLLMs) offer promising capabilities for visual understanding, existing MLLM-based approaches struggle to precisely localize infrared small targets. In this paper, we propose IRSTD-Agent, an agentic framework for infrared small target detection through dynamic visual search. The framework enables an MLLM to adaptively determine where and at what scale to inspect an image and progressively gather fine-grained visual evidence for precise target localization. Five complementary visual tools (PROPOSAL, ZOOM, DETECT, DROP and REFINE) support object candidate discovery, adaptive observation, target localization, hypothesis rejection, and target extent refinement, together enabling a coordinated search process over original-resolution images. To teach the MLLMs to conduct this search, we introduce Zoom-guided Interaction Learning, which uses annotation-derived interaction trajectories to supervise tool selection and the corresponding arguments. Through extensive experiments on WideIRSTD-Full and IRSTD-1k datasets, we demonstrate that IRSTD-Agent outperforms the evaluated vision-language models and enhances the precise localization capabilities of MLLMs in IRSTD tasks.