🤖 AI Summary
This study addresses the lack of a standardized protocol for inference window configuration—discrete versus overlapping—in reconstruction-based time series anomaly detection, which hinders fair performance comparison across methods. The authors systematically investigate the impact of inference stride on detection performance and propose a unified evaluation framework encompassing consistent training, hyperparameter tuning, and multi-random-seed assessment. Evaluated on the TSB-AD and UCR benchmarks across diverse reconstruction models—including PCA, DLinear, autoencoders, TimesNet, and Transformers—the analysis reveals, for the first time, that overlapping windows substantially enhance detection performance and alter method rankings. On average, overlapping windows yield a 28% relative performance improvement across all models, demonstrating their critical role in reliable evaluation. Moreover, reconstruction-based baselines exhibit strong and practical performance on both benchmarks, underscoring the need for reproducible evaluation standards in the field.
📝 Abstract
Reconstruction-based methods are widely used for time series anomaly detection, where models are trained to reconstruct subsequences, and anomalies are identified through reconstruction errors. However, reported results are often hard to compare due to heterogeneous evaluation practices and underspecified inference procedures. In this paper, we revisit reconstruction-based anomaly detection in the univariate offline setting and study the role of the inference stride, which controls whether subsequences are processed as disjoint windows or with overlap. We propose a unified training, tuning, and multi-seed evaluation protocol on the curated TSB-AD benchmark, and study how overlapping inference affects anomaly detection performance for a range of reconstruction models, including PCA-based baselines, DLinear, an AutoEncoder, TimesNet, and Transformer variants. The results show that across all models, overlapping windows yield consistent improvements, with average relative gain up to +28%, and can alter method rankings. We further analyze variability across datasets, random seeds, and hyperparameter configurations. Finally, we complement the benchmark study with an evaluation on the full UCR archive using localization criteria aligned with sliding-window reconstruction. Overall, our results highlight that reconstruction-based anomaly detection performance depends not only on model architecture and training, but also on inference choices, motivating a clear and reproducible protocol. Our results show that reconstructionbased baselines achieve strong performance on both TSB-AD and UCR benchmarks, supporting them as competitive and practical approaches for univariate time series anomaly detection.