Revisiting OmniAnomaly for Anomaly Detection: performance metrics and comparison with PCA-based models

📅 2026-03-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of fairly comparing deep learning and classical statistical methods in multivariate time series anomaly detection (MTSAD), where inconsistent evaluation protocols often obscure true performance differences. Under a unified thresholding strategy and evaluation pipeline, the authors conduct 100 repeated experiments on each of the 28 machines in the SMD dataset, comparing OmniAnomaly—a stochastic recurrent neural network-based model—against a principal component analysis (PCA) baseline. Without point-adjustment strategies, PCA consistently matches or outperforms OmniAnomaly, with substantial performance variation observed across different machines. These findings suggest that the purported advantages of complex models in current MTSAD benchmarks may be overstated, highlighting the critical influence of evaluation protocols on reported results. The work provides a rigorously reproducible framework for future comparative studies in this domain.

Technology Category

Machine Learning: Evaluation and AnalysisData Mining & Knowledge Management: Anomaly/Outlier DetectionNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP Models

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methodsGraph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphs
📝 Abstract
Deep learning models have become the dominant approach for multivariate time series anomaly detection (MTSAD), often reporting substantial performance improvements over classical statistical methods. However, these gains are frequently evaluated under heterogeneous thresholding strategies and evaluation protocols, making fair comparisons difficult. This work revisits OmniAnomaly, a widely used stochastic recurrent model for MTSAD, and systematically compares it with a simple linear baseline based on Principal Component Analysis (PCA) on the Server Machine Dataset (SMD). Both methods are evaluated under identical thresholding and evaluation procedures, with experiments repeated across 100 runs for each of the 28 machines in the dataset. Performance is evaluated using Precision, Recall and F1-score at point-level, with and without point-adjustment, and under different aggregation strategies across machines and runs, with the corresponding standard deviations also reported. The results show large variability across machines and show that PCA can achieve performance comparable to OmniAnomaly, and even outperform it when point-adjustment is not applied. These findings question the added value of more complex architectures under current benchmarking practices and highlight the critical role of evaluation methodology in MTSAD research.
Problem

Research questions and friction points this paper is trying to address.

multivariate time series anomaly detection
evaluation protocol
thresholding strategy
fair comparison
performance metrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

OmniAnomaly
PCA
evaluation methodology
multivariate time series anomaly detection
point-adjustment
🔎 Similar Papers
No similar papers found.