Score
Designs and implements detectors that identify and localize anomalous inputs by training denoising autoencoders (DAEs) on nominal data and using reconstruction error as a per-window or per-sample anomaly score. This includes choosing corruption/denoising strategies, computing and aggregating reconstruction-based anomaly scores, setting thresholds, and mapping high-error regions to anomalous segments.
Anomaly detection (AD) in high-dimensional, unstructured data faces persistent challenges in model expressiveness and interpretability. Method: This paper presents a systematic survey of over 180 deep learning–based AD studies published between 2018 and 2024, unifying reconstruction-based (e.g., autoencoders, GANs) and prediction-based (e.g., LSTMs/Transformers, GNNs) paradigms for the first time. It proposes a hybrid framework that jointly optimizes interpretability and performance by integrating statistical hypothesis testing, ensemble learning, and deep models. A multimodal taxonomy is constructed, and extensive evaluation is conducted across benchmarks including UCR and KDD Cup. Contribution/Results: The framework achieves an average 12.3% improvement in F1-score. The authors publicly release an evaluation matrix and practical implementation guidelines, and identify six open challenges and future research directions.
This paper challenges the foundational assumption of autoencoder-based anomaly detection—that anomalous samples incur higher reconstruction error—demonstrating its theoretical invalidity. Method: Through rigorous theoretical analysis and cross-modal empirical evaluation on tabular data and real-world images, we prove that anomalies lying far from the normal manifold can be reconstructed with zero error by both linear and nonlinear autoencoders. Such hazardous extrapolation behavior causes critical misclassification of anomalies as normal, severely compromising reliability in safety-critical applications. Contributions: (1) We derive tight theoretical bounds on reconstruction error, clarifying the failure mechanism; (2) We systematically validate zero-error reconstruction across multiple real-world datasets spanning diverse anomaly types; (3) We expose the fundamental limitations of prevailing autoencoder methods under high-assurance requirements, providing a crucial foundation for rethinking trustworthy anomaly detection.
In unsupervised anomaly detection, autoencoders (AEs) often suffer from over-generalization to anomalous samples, resulting in underestimated reconstruction errors and high false-negative rates. To address this, we propose a decoupled joint-training framework that disentangles and fuses AE latent-space representations with reconstruction-quality features, yielding more discriminative composite features. Furthermore, we introduce an optimizable contrastive Gaussian noise distribution to enhance the robustness of noise-contrastive estimation (NCE)-based density modeling. Evaluated on multiple standard benchmarks, our method achieves state-of-the-art performance among AE-based approaches—significantly reducing false-negative rates while maintaining low false-positive rates. This work establishes a novel paradigm for improving the reliability of AE-based models in anomaly detection.
Unsupervised industrial anomaly detection faces the challenge that single-pass reconstruction struggles to simultaneously suppress anomalies and preserve fine-grained details. To address this, we propose a recursive autoencoder framework that iteratively refines reconstruction for precise anomaly localization. Specifically, we design a Cross-Recursive Detection (CRD) module to model the dynamic evolution of anomalies across recursion steps, introduce a Detail-Preserving Network (DPN) to recover high-frequency textures, and enforce cross-recursive consistency constraints alongside an unsupervised anomaly scoring mechanism. Without relying on diffusion processes, our method achieves performance competitive with state-of-the-art diffusion-based models—while using only 10% of their parameters and significantly accelerating inference. Extensive experiments demonstrate substantial improvements over non-diffusion-based SOTA methods on benchmark industrial datasets, including MVTec-AD.
Existing anomaly detection models suffer from training bias and degraded generalization due to the lack of a unified standard for anomaly modeling. This paper addresses the limitation of reconstruction-based methods—specifically, their failure to account for inter-class discrepancies in anomaly characteristics during data augmentation—by proposing the first composable augmentation framework tailored for reconstruction models. We first systematically identify key factors by which synthetic anomalies influence reconstruction training. Then, we design a class-aware augmentation composition mechanism coupled with a decoupled, multi-stage training strategy comprising feature-space perturbation, class-conditional augmentation selection, and parameter freezing. On MVTec-AD, our method significantly outperforms state-of-the-art approaches, especially in object-level anomaly detection. Moreover, on a newly constructed multi-characteristic synthetic anomaly benchmark, it demonstrates superior cross-class generalization capability.
Existing generative anomaly detection methods suffer from insufficient reconstruction quality, limiting detection accuracy and efficiency in industrial applications. To address this, we propose a novel “noise-to-norm” reconstruction paradigm: a one-step norm-guided diffusion model is designed with image norm as the explicit reconstruction objective, enabling high-fidelity anomaly-free reconstruction. We further introduce a multi-scale noise fusion mechanism and a reconstruction-preserving fast denoising strategy, accelerating inference by up to two orders of magnitude. Additionally, we devise a dual-branch architecture integrating reconstruction and segmentation, coupled with pixel-wise similarity analysis to generate precise anomaly score maps. Our method achieves state-of-the-art performance across four standard benchmarks, with substantially improved reconstruction fidelity and inference speed comparable to conventional (non-generative) approaches. The source code is publicly available.
This work addresses the challenge that existing unsupervised methods struggle to reliably detect subtle and noisy anomalies in complex time series, often being misled by noise in normal samples and missing near-normal anomalies. To overcome this limitation, we propose a novel unsupervised anomaly detection framework that integrates active learning: it enhances temporal dependency modeling through a masked time series reconstruction feedback mechanism and employs a minimax optimization strategy to differentially treat normal and anomalous samples, thereby improving robustness against noise and weak anomalies. Extensive experiments across four multivariate time series datasets and seven backbone models demonstrate that our method achieves an average AUC improvement of 12.39%, significantly outperforming current unsupervised approaches.
Detecting anomalies in images and video is an essential task for multiple real-world problems, including industrial inspection, computer-assisted diagnosis, and environmental monitoring. Anomaly detection is typically formulated as a one-class classification problem, where the training data consists solely of nominal values, leaving methods built on this assumption susceptible to training label noise. We present a dataset folding method that transforms an arbitrary one-class classifier-based anomaly detector into a fully unsupervised method. This is achieved by making a set of key weak assumptions: that anomalies are uncommon in the training dataset and generally heterogeneous. These assumptions enable us to utilize multiple independently trained instances of a one-class classifier to filter the training dataset for anomalies. This transformation requires no modifications to the underlying anomaly detector; the only changes are algorithmically selected data subsets used for training. We demonstrate that our method can transform a wide variety of one-class classifier anomaly detectors for both images and videos into unsupervised ones. Our method creates the first unsupervised logical anomaly detectors by transforming existing methods. We also demonstrate that our method achieves state-of-the-art performance for unsupervised anomaly detection on the MVTec AD, ViSA, and MVTec Loco AD datasets. As improvements to one-class classifiers are made, our method directly transfers those improvements to the unsupervised domain, linking the domains.
This work addresses the overreliance on complex architectures in time series anomaly detection by proposing JuRe (Just Repair), an exceptionally minimalist approach that employs only a single depthwise-separable convolutional residual block to construct a denoising network. During training, JuRe learns to reconstruct corrupted temporal windows, while at inference time it leverages a parameter-free structural discrepancy function for scoring anomalies. Notably, JuRe eschews attention mechanisms, latent variables, and adversarial components, demonstrating for the first time that—under proper manifold projection—the design of the training objective is more decisive than model capacity for performance. On the TSB-AD multivariate benchmark, JuRe achieves an AUC-PR of 0.404 (ranking second), and on the UCR univariate datasets, it attains an AUC-PR of 0.198 (also second), outperforming all neural baselines in both AUC-PR and VUS-PR metrics.
This work addresses the instability and insufficient sensitivity of denoising score matching (DSM) in tabular anomaly detection, particularly under scenarios lacking validation sets or label information, where selecting an appropriate perturbation scale is challenging. To overcome this, the authors propose K-DSM, a method that adaptively assigns a noise level to each feature based on its kurtosis, enabling efficient single-scale DSM training. Additionally, an exponential moving average (EMA) teacher filtering mechanism is introduced to mitigate data contamination. The framework eliminates the need for multi-scale or noise-conditioned training, substantially reducing hyperparameter dependence while enhancing coverage in low-density regions and discrimination accuracy in high-density regions. Experimental results demonstrate that K-DSM achieves state-of-the-art performance under semi-supervised settings and remains highly effective even in fully unsupervised scenarios with contaminated data.
To address overfitting in zero-shot anomaly generation caused by the absence of authentic anomaly samples—rendering conventional fine-tuning ineffective—this paper proposes a novel diffusion-based paradigm that requires neither training nor access to any anomaly samples. Methodologically, we design a dual-branch contrastive denoising architecture: subtle prompt perturbations induce divergent denoising trajectories, and accumulated denoising residuals enable interpretable, pixel-level anomaly localization. We further enhance generation stability and local controllability via token-level prompt refinement, latent-space localized inpainting, and constrained spatial attention biasing. To our knowledge, this is the first approach achieving high-fidelity, explainable, *fully* zero-shot localized anomaly synthesis—without any anomaly annotations or model adaptation. Extensive experiments on multiple public benchmarks demonstrate significant improvements in downstream anomaly detection performance.