Score
Designs and implements metrics, taxonomies, detection signals, and analytical frameworks that describe how the underlying data-generating process or label concept changes over time—including type (abrupt, gradual, recurring), magnitude, duration, and theoretical properties. Produces empirical and theoretical characterizations that quantify drift and relate those drift attributes to learner performance and behavior.
This study addresses the absence of a unified evaluation framework for concept drift detection, where metrics such as classification accuracy are frequently misapplied and fail to faithfully reflect detection performance. For the first time, it systematically links eight categories of drift detection quality metrics to classifier performance through extensive experiments on seven synthetic non-stationary data streams, while explicitly modeling the dynamic characteristics of drift. The work reveals the inherent limitations of classification accuracy in evaluating drift detection and identifies a more informative combination of metrics. These findings provide both empirical evidence and theoretical grounding for establishing a standardized and reliable evaluation methodology in the field of concept drift detection.
This paper addresses the critical problem of concept drift detection in supervised learning. We propose a multivariate exponentially weighted moving average (MEWMA) control chart method based on shifts in the mean of Fisher score vectors. Our key contributions are: (1) integrating nested bootstrapping with the 0.632+ variance correction to accurately estimate control limits—thereby significantly improving false alarm rate (FAR) control, especially under small-sample and stringent FAR constraints; and (2) eliminating the need for an initial labeled dataset to calibrate control limits, enabling full utilization of all available samples for model training and thus jointly enhancing detection sensitivity and model accuracy. Empirical evaluation demonstrates that our method achieves superior FAR control precision compared to conventional approaches, with particularly pronounced advantages in small-sample regimes and under strict false alarm constraints.
This study addresses the performance degradation of machine learning models caused by concept drift in dynamic data streams. It systematically analyzes the characteristics of concept drift and theoretically investigates, alongside empirical evaluation, the behavior of multiple learner-based detection algorithms under diverse drift scenarios—including abrupt and gradual shifts. Through comprehensive experiments on both synthetic and real-world datasets, the work compares the behavioral patterns and applicability of various detection methods, thereby deepening the understanding of underlying drift mechanisms. The findings elucidate the relative strengths and limitations of different detectors across heterogeneous environments, offering robust empirical guidance for algorithm selection in practical applications.
This work addresses the challenge of concept drift in data streams, which often degrades model performance, by proposing FiCSUM—a novel framework that effectively distinguishes between emerging and recurring concepts. FiCSUM is the first to integrate supervised and unsupervised multidimensional meta-information features to construct highly discriminative “concept fingerprints.” It further incorporates a dynamic weighting mechanism that adaptively identifies concept changes. The approach operates through meta-feature extraction, dynamic weighting, fingerprint vector construction, and similarity-based detection, significantly enhancing the ability to detect both new and reappearing concepts. Extensive experiments on 11 real-world and synthetic datasets demonstrate that FiCSUM consistently outperforms state-of-the-art methods in both detection accuracy and concept modeling fidelity.
In dynamic time series, the co-occurrence of concept drift and anomalies impedes their discrimination, leading to false detections, error propagation, and excessive model updates. This paper proposes AnDri—the first unified framework for joint concept drift detection and anomalous subsequence identification. Its key contributions are: (1) Adjacency-aware Hierarchical Clustering (AHC), which preserves temporal locality to precisely segment evolving normal patterns; (2) an adaptive normal-pattern update mechanism integrated with sliding-window subsequence modeling; and (3) a joint drift–anomaly discrimination strategy. Evaluated on diverse real-world time series datasets, AnDri achieves an average 12.7% improvement in F1-score and reduces spurious model update rate by 63%. By mitigating error accumulation, it significantly enhances robustness for downstream time-series analysis tasks.
Existing concept drift methodologies struggle to handle the complex, multidimensional, and non-stationary dynamics inherent in data streams encountered by autonomous learning systems. This work proposes a unified three-dimensional classification framework that systematically characterizes drift phenomena across temporal, data, and model streams, thereby integrating paradigms such as drift adaptation, continual learning, and temporal generalization. Through a comprehensive systematic review of 193 studies, the paper introduces and formally delineates novel concepts—including representational drift, semantic drift, and policy instability—exposes critical limitations of current approaches, identifies fundamental open challenges, and offers a clear roadmap toward building intelligent learning systems capable of sustainable, autonomous evolution.
This study addresses the performance degradation of malware classification models caused by concept drift due to the continuous evolution of malicious software. To mitigate this issue, the authors propose an adaptive retraining mechanism that updates the model only when a distribution shift is detected, thereby balancing classification accuracy and training efficiency. The approach innovatively employs One-Class Support Vector Machine (OCSVM) for concept drift detection and compares its effectiveness against Minibatch K-Means and Maximum Mean Discrepancy (MMD). Experimental results demonstrate that OCSVM achieves classification accuracy comparable to periodic retraining while substantially reducing the number of retraining events. Overall, the proposed method offers a superior Pareto trade-off between accuracy and computational overhead compared to baseline approaches.
Concept drift detection is widely employed in data stream learning, yet its efficacy remains inadequately validated, and it often fails to distinguish genuine distributional shifts from spurious drifts induced by the detection mechanism itself. This work introduces the notion of the “window dilemma,” revealing that sliding-window–based detection is fundamentally ill-posed: observed drift may stem from windowing choices rather than actual changes in the underlying data-generating process. Through theoretical analysis, illustrative examples, and large-scale empirical comparisons, the study systematically evaluates a range of drift detectors against non-drift-aware adaptive and batch learning methods. The results demonstrate that conventional batch learners consistently outperform drift-detection–based streaming classifiers across most scenarios, thereby raising fundamental questions about the necessity and practical utility of prevailing concept drift detection paradigms.
Existing drift detection methods rely on fixed thresholds, struggling to simultaneously minimize false positives and false negatives while exhibiting poor robustness to distributional shifts. This paper proposes a dynamic thresholding mechanism, establishing—through segmented performance analysis and rigorous theoretical proof—that time-varying thresholds are statistically superior to any fixed threshold. The resulting adaptive algorithm requires no prior knowledge and optimizes thresholds online to balance detection sensitivity and stability. Furthermore, a comparative-phase module is introduced to integrate seamlessly with mainstream detectors (e.g., ADWIN, DDM) and extend applicability to multimodal real-world data—including images and tabular datasets. Experiments across 12 synthetic and real-world benchmarks demonstrate that our method reduces false positive rate by 37.2% on average and improves drift recall by 29.5%, significantly enhancing model performance retention under continual learning.
This study addresses the lack of a unified and fair evaluation benchmark for concept drift detection methods, which hinders meaningful cross-method comparisons. To this end, the authors propose a systematic evaluation framework that injects controlled, diverse types of drift—such as class prior changes and label swaps—into seven real-world datasets via Monte Carlo simulation. The framework introduces time-sensitive metrics, including F1 detection score and normalized detection delay, and employs a leave-one-dataset-out hyperparameter optimization strategy to enhance generalization. Through comprehensive evaluation of 14 state-of-the-art methods under this framework, the work establishes the first performance benchmarks for both abrupt and gradual drift scenarios, revealing the relative strengths, weaknesses, and applicability conditions of existing approaches.