🤖 AI Summary
This study addresses the limitations of existing streaming time-series anomaly detection methods, which are largely derived from outlier detection, neglect temporal characteristics, and lack validation in large-scale real-world scenarios. We establish a unified evaluation framework that compares the accuracy and efficiency of streaming versus static methods online following an initial training phase. Furthermore, we introduce TSB-drift, a real-world distribution drift dataset designed to isolate applicable scenarios. Through the first large-scale empirical study, we reveal the counterintuitive finding that static methods significantly outperform streaming approaches in most streaming settings. This work exposes fundamental design flaws in current streaming algorithms and calls for a rethinking of the integration paradigms used to incorporate streaming capabilities into time-series anomaly detection systems.
📝 Abstract
Time series anomaly detection (TSAD) is increasingly deployed in streaming settings, where data arrive sequentially and may exhibit non-stationarity. As a result, several works from the recent literature propose streaming anomaly detection methods that rely on incremental updates to adapt over time. However, most of these approaches originate from the streaming outlier detection literature and largely ignore core characteristics of time series anomalies. Moreover, their empirical evaluation is typically conducted on synthetic or small-scale benchmarks with limited diversity, making it unclear whether streaming methods are truly advantageous in realistic TSAD scenarios. In this work, we carry out the first large-scale experimental study comparing streaming and static TSAD methods under a unified streaming evaluation benchmark. We consider a realistic setting in which an initial batch of data is available for model training, followed by online evaluation of both detection accuracy and computational efficiency. In addition, we propose a distribution-drift dataset of real time series, called TSB- drift, to isolate scenarios where streaming updates are theoretically justified. Our results show that, contrary to common assumptions, static TSAD methods significantly outperform streaming approaches in most streaming settings. Such finding highlights a critical gap between the design of existing streaming methods and the requirements of modern TSAD, and calls for a rethinking of how streaming capabilities should be integrated into TSAD.