GRC-Net: Global Representation Consistency Network for Unsupervised Multimodal Anomaly Detection

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing local patch-based anomaly detection methods are susceptible to reconstruction noise, resulting in unstable errors within normal regions and hindering the effective identification of multimodal structural anomalies. To address this, this work proposes GRC-Net, which introduces a global representation consistency mechanism. By leveraging global tokens and attention MLPs to capture holistic context, the method suppresses reconstruction noise and integrates a stable reconstruction module for unsupervised multimodal anomaly detection, fundamentally resolving the reconstruction instability inherent in local representations. Extensive experiments on the MVTec 3D-AD and Eyecandies datasets demonstrate that GRC-Net significantly outperforms existing methods in both image-level and pixel-level detection performance, substantially enhancing overall detection robustness.
📝 Abstract
Automated quality inspection is essential for ensuring product reliability in manufacturing.While image-based methods effectively capture appearance-related defects, these methods are limited in detecting structural and geometric anomalies, motivating multimodal approaches incorporating 3D information. However, existing methods mainly rely on local patch-level representations, which often lead to unstable reconstruction errors even in normal regions. To address this limitation, we propose GRC-Net, which integrates a global-attention MLP to enforce global representation consistency across patch embeddings with a stable reconstruction module to improve reconstruction stability. The proposed method captures holistic contextual information through a global token and suppresses reconstruction noise by minimizing discrepancies between original and predicted embeddings. Experiments on MVTec 3D-AD and Eyecandies demonstrate that GRC-Net consistently outperforms existing methods at both image and pixel levels. Qualitative results further demonstrate reduced reconstruction errors in normal regions and more distinct reconstruction differences between normal and anomalous regions.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Anomaly Detection
Unsupervised Learning
Quality Inspection
Reconstruction Stability
3D-AD
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unsupervised Multimodal Anomaly Detection
Global Representation Consistency
Global-Attention MLP
Reconstruction Stability
Patch Embeddings
🔎 Similar Papers
2024-05-29arXiv.orgCitations: 0
💼 Related Jobs
No related jobs found.
S
Seyoung Jeong
Jeonbuk National University, Jeonju, Republic of Korea
J
Jong Pil Yun
Korea Institute of Industrial Technology (KITECH), Incheon, Republic of Korea; Chung-Ang University, Seoul, Republic of Korea
Sang Jun Lee
Sang Jun Lee
stradvision
SLAMSensor FusionNeural RenderingSpatial AI