🤖 AI Summary
This study addresses the limitations of multimodal large language models in fine-grained perception and single-model performance for industrial anomaly detection and reasoning by proposing the SiGMA framework. SiGMA introduces a novel spatially-guided multi-agent architecture that coordinates heterogeneous agents with visual expert modules, incorporates a searcher for knowledge retrieval, and employs an unsupervised reliability controller to dynamically weigh the quality of multi-source evidence while supporting seamless integration of new models. Evaluated on the MMAD benchmark, the proposed method achieves an average accuracy of 85.2%, surpassing the strongest baseline by 4.0% and approaching human-level performance. Notably, deploying only three agents with 9B parameters yields an accuracy of 84.4%.
📝 Abstract
A single multimodal large language model (MLLM) struggles to excel simultaneously at detection, localization, description, and reasoning in multimodal industrial anomaly understanding (MM-IAU). We show that in-domain training does not close this gap. On MMAD, a widely adopted MM-IAU benchmark, trained specialists reach at most 75.5% accuracy in defect localization, against 92.3% for human experts, and even detect anomalies less accurately than their untrained base model. Meanwhile, different MLLMs offer complementary strengths but share this weakness in fine-grained perception, so combining them alone cannot remove it. We therefore propose SiGMA, a spatially grounded multi-agent framework that divides labor between heterogeneous MLLM agents and a dedicated visual defect expert. A multimodal searcher supplies industrial knowledge and normal references, the defect expert turns query-reference comparison into calibrated anomaly evidence, and a label-free reliability controller weighs each source by task-wise competence and query-level evidence quality. SiGMA reaches 85.2% average accuracy on MMAD, 4.0% above the strongest trained specialist and Gemini-2.5-Pro and within 1.5% of human experts. Even with three agents of at most 9B parameters, it reaches 84.4%, and new MLLMs join without retraining.