Augmenting Visual Anomaly Detection with Automated Interpretability

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of false alarms in visual anomaly detection caused by the entanglement of genuine anomalies with benign variation signals. We propose a novel framework integrating automated interpretability with feature-level intervention. Specifically, residuals extracted via PatchCore are decomposed into independent features using sparse autoencoders (SAEs) and automatically annotated by multimodal large language models (MLLMs). This enables precise suppression of benign interference and targeted amplification of anomalous signals through reconstructed embeddings for optimized detection. The proposed method achieves accurate decoupling of anomalous signals, yielding significant improvements in AUROC across multiple benchmarks. Notably, it demonstrates superior robustness under challenging conditions involving data corruption and acquisition shifts.
📝 Abstract
Visual anomaly detectors identify deviations from known-normal data, but their anomaly signals may mix evidence of actual anomalies with benign visual variation. We investigate whether automated interpretability can augment visual anomaly detectors by identifying and intervening on different components of this signal. We decompose PatchCore nearest-normal residuals into sparse features using Sparse Autoencoders (SAEs), and provide high-activation and contrastive non-active examples to a Multimodal LLM, which describes each feature and labels it as anomaly, distractor, or uncertain. These labels guide interventions in the SAE hidden representation, where distractor features are suppressed and anomaly features amplified. The edited representation is then used to reconstruct patch embeddings, which are rescored with PatchCore. Across 40 categories from four benchmarks, applying both interventions jointly improves macro-average image-level AUROC from 0.8724 to 0.8857 on source data and from 0.8066 to 0.8210 under synthetic corruptions. On three additional RobustAD categories with real acquisition shifts, the same interventions improve AUROC from 0.8745 to 0.9056 on source data and from 0.6069 to 0.6599 under real acquisition shifts. Finally, individual feature interventions across all 43 categories show that the MLLM labels are aligned in aggregate with how features differently affect normal and anomalous images.
Problem

Research questions and friction points this paper is trying to address.

Visual Anomaly Detection
Automated Interpretability
Sparse Autoencoders
Robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Anomaly Detection
Sparse Autoencoders
Automated Interpretability
Multimodal LLM
Representation Intervention
🔎 Similar Papers
No similar papers found.