Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of multimodal large language models in fine-grained perception and single-model performance for industrial anomaly detection and reasoning by proposing the SiGMA framework. SiGMA introduces a novel spatially-guided multi-agent architecture that coordinates heterogeneous agents with visual expert modules, incorporates a searcher for knowledge retrieval, and employs an unsupervised reliability controller to dynamically weigh the quality of multi-source evidence while supporting seamless integration of new models. Evaluated on the MMAD benchmark, the proposed method achieves an average accuracy of 85.2%, surpassing the strongest baseline by 4.0% and approaching human-level performance. Notably, deploying only three agents with 9B parameters yields an accuracy of 84.4%.
📝 Abstract
A single multimodal large language model (MLLM) struggles to excel simultaneously at detection, localization, description, and reasoning in multimodal industrial anomaly understanding (MM-IAU). We show that in-domain training does not close this gap. On MMAD, a widely adopted MM-IAU benchmark, trained specialists reach at most 75.5% accuracy in defect localization, against 92.3% for human experts, and even detect anomalies less accurately than their untrained base model. Meanwhile, different MLLMs offer complementary strengths but share this weakness in fine-grained perception, so combining them alone cannot remove it. We therefore propose SiGMA, a spatially grounded multi-agent framework that divides labor between heterogeneous MLLM agents and a dedicated visual defect expert. A multimodal searcher supplies industrial knowledge and normal references, the defect expert turns query-reference comparison into calibrated anomaly evidence, and a label-free reliability controller weighs each source by task-wise competence and query-level evidence quality. SiGMA reaches 85.2% average accuracy on MMAD, 4.0% above the strongest trained specialist and Gemini-2.5-Pro and within 1.5% of human experts. Even with three agents of at most 9B parameters, it reaches 84.4%, and new MLLMs join without retraining.
Problem

Research questions and friction points this paper is trying to address.

multimodal industrial anomaly understanding
fine-grained perception
defect localization
multimodal large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Framework
Multimodal Large Language Models
Industrial Anomaly Understanding
Spatial Grounding
Reliability Controller
🔎 Similar Papers
No similar papers found.
X
Xingwu Zhang
Hunan University
D
Duanyang Du
University of Aberdeen
Huiling Zhu
Huiling Zhu
Unknown affiliation
Wireless/Mobile Communications
J
Jiayue Dai
University of Aberdeen
Y
Yixiao Liu
Hunan University
G
Guozhi Liu
South China University of Technology
Zhihan Zhang
Zhihan Zhang
PhD student, University of Notre Dame
Natural Language Processing
Z
Zijun Long
Hunan University