Score
Designs, implements, and evaluates algorithms and statistical procedures that detect, localize, and quantify distributional or structural changes in temporal or streaming data, producing low‑latency change‑point estimates, alarms, or concept‑drift signals. This work covers online/sequential and batch methods, unsupervised and supervised detectors, temporal/parameter drift analysis, and tools to verify, summarize, and trigger retraining or adaptation when change events are detected.
Existing data stream learning research often relies on unrealistic assumptions—such as single-pass processing and strict online constraints—leading to ill-defined problem formulations, biased evaluation protocols, and misalignment with industrial requirements. Method: This paper systematically critiques and deconstructs these restrictive assumptions, proposing a “de-paradigmized” framework that centers modeling on concept drift and temporal dependence while abandoning rigid formal stream constraints; algorithmic design integrates time-series analysis, concept drift detection, robust statistical learning, and privacy-preserving techniques—rejecting isolated development of bespoke streaming algorithms. Contribution/Results: The work yields a methodology guide grounded in industrial practice, fostering renewed consensus between academia and industry. It significantly enhances model robustness, interpretability, and privacy compliance in real-world dynamic environments.
This work addresses the challenge of accurately attributing detected change points in multivariate time series to specific subsets of variables. The authors propose a post-hoc, nonparametric testing framework that, after an offline change point has been identified, determines whether the change occurs in one of two pre-specified coordinate blocks or in both. Built upon two-sample nonparametric hypothesis testing, the method offers rigorous theoretical guarantees for Type I error control. Empirical evaluations on both synthetic and real-world datasets demonstrate that the proposed approach achieves high attribution accuracy and strong robustness in identifying the components responsible for the change.
This study addresses the challenge of detecting and localizing mean change points in high-dimensional dependent time series, particularly under non-Gaussian distributions and temporal dependence. The authors propose an adaptive method that integrates a quadratic-form CUSUM statistic with a coordinate-wise maximum statistic to effectively capture dense and sparse changes, respectively, while employing a weighting scheme to distinguish interior from boundary change points. A key theoretical contribution lies in establishing the limiting distributions and asymptotic independence of these two statistics under non-Gaussian dependence, thereby providing rigorous justification for a Cauchy combination test. Coupled with wild binary segmentation, the approach achieves consistent estimation of multiple change points. Theoretical analysis confirms the validity of centering and scaling procedures, and extensive numerical experiments demonstrate superior detection accuracy and localization efficiency across diverse scenarios.
This study addresses the performance degradation of machine learning models caused by concept drift in dynamic data streams. It systematically analyzes the characteristics of concept drift and theoretically investigates, alongside empirical evaluation, the behavior of multiple learner-based detection algorithms under diverse drift scenarios—including abrupt and gradual shifts. Through comprehensive experiments on both synthetic and real-world datasets, the work compares the behavioral patterns and applicability of various detection methods, thereby deepening the understanding of underlying drift mechanisms. The findings elucidate the relative strengths and limitations of different detectors across heterogeneous environments, offering robust empirical guidance for algorithm selection in practical applications.
This work addresses the problem of efficient online change-point detection in both univariate and multivariate data streams by introducing a novel method grounded in the Focus algorithm family. Leveraging the generalized likelihood ratio test, the approach enables exact detection of a single change point without requiring approximations. By exploiting the relationship between candidate change-point locations and the geometric structure of the data, it achieves a computational complexity of approximately $\log(n)^d$ per iteration. Notably, this is the first method to support exponential-family models, nonparametric settings, and autoregressive data under no approximation assumptions, integrating natural exponential-family modeling, empirical cumulative distribution functions, and geometric optimization techniques. The accompanying R/Python software package substantially enhances the efficiency and applicability of change-point detection in high-dimensional streaming data.
This work addresses the limitations of existing online change-point detection methods, which typically assume independent and identically distributed (i.i.d.) data and thus struggle with real-world time series exhibiting autocorrelation, often resulting in high false alarm rates or detection delays. To overcome this, the authors extend the generalized likelihood ratio (GLR) test to p-th order autoregressive (AR(p)) processes and integrate it with an enhanced FOCUS algorithm, yielding a computationally efficient online detector termed AR(p)-FOCUS. This approach is the first to explicitly model temporal dependencies within the GLR framework, achieving significantly improved detection performance while maintaining an average per-iteration computational complexity of O(log n). Empirical evaluations demonstrate that AR(p)-FOCUS outperforms conventional i.i.d.-based methods on autocorrelated data and exhibits strong effectiveness and practicality on real-world telecommunications datasets.
This study addresses the challenge of rapid change-point detection in high-dimensional multisensor systems under structural constraints and limited sensing resources. By integrating sparse modeling, heterogeneous data fusion, and a resource-adaptive sequential sampling strategy, the work extends classical change-point detection theory to large-scale, resource-constrained sensing scenarios and incorporates machine learning to handle cases with unknown system models. The proposed approach unifies sparse signal processing, multi-stream statistical decision-making, and resource-constrained optimization to enable simultaneous detection of multiple change points. This framework significantly enhances both applicability and scalability in high-dimensional, heterogeneous, and resource-limited environments while maintaining high detection efficiency.