🤖 AI Summary
This study addresses the inefficiency of visual iteration and the susceptibility to local evidence loss in multimodal models for industrial fine-grained defect detection. To this end, we propose Anomaly-LR, a framework that leverages visual latent space representations and a global-to-local progressive refinement strategy, integrated with multimodal large language models to enable defect-oriented reasoning. Furthermore, we introduce IAD-LR-22K, the first instruction dataset tailored for latent reasoning in industrial anomaly detection, which effectively preserves local defect evidence without relying on external tools. Experimental results demonstrate that the proposed framework achieves state-of-the-art performance across multiple benchmarks, significantly outperforming existing methods of comparable scale.
📝 Abstract
Industrial anomaly detection (IAD) is evolving beyond conventional detection and localization toward multimodal inspection systems that can describe, explain, and reason about fine-grained defects. Although recent multimodal large language model (MLLM)-based methods improve anomaly understanding through textual reasoning and visual guidance, they face two limitations in fine-grained inspection. First, their visual refinement often requires iteratively revisiting local image regions or augmenting with additional tools. Second, the resulting local defect evidence may not be reliably preserved throughout subsequent reasoning. To address these, we propose Anomaly-LR, a defect-grounded latent reasoning framework that first forms a global understanding of the input and then progressively refines anomaly-relevant representations directly in the visual latent space. We further construct IAD-LR-22K, the first IAD instruction dataset designed for latent reasoning, containing 22,228 image-question instances from 4,523 industrial images, with global textual reasoning traces and region-level visual annotations. Extensive experiments show that Anomaly-LR achieves state-of-the-art performance among comparable-scale methods across multiple IAD benchmarks, without requiring external references or tools. The code and data will be released at https://github.com/Yen666/Anomaly-LR.