Score
Designs, implements, and evaluates algorithms, mechanisms, and end-to-end pipelines that update models or model components incrementally from streaming data so the system adapts continuously without retraining from scratch. This includes developing and analyzing online learning algorithms and theory, online model-adaptation methods and systems, and streaming spectral or matrix-update techniques such as incremental/streaming SVD with temporal forgetting or reweighting to maintain or modify bases and preconditioners over time.
Existing data stream learning research often relies on unrealistic assumptions—such as single-pass processing and strict online constraints—leading to ill-defined problem formulations, biased evaluation protocols, and misalignment with industrial requirements. Method: This paper systematically critiques and deconstructs these restrictive assumptions, proposing a “de-paradigmized” framework that centers modeling on concept drift and temporal dependence while abandoning rigid formal stream constraints; algorithmic design integrates time-series analysis, concept drift detection, robust statistical learning, and privacy-preserving techniques—rejecting isolated development of bespoke streaming algorithms. Contribution/Results: The work yields a methodology guide grounded in industrial practice, fostering renewed consensus between academia and industry. It significantly enhances model robustness, interpretability, and privacy compliance in real-world dynamic environments.
This paper addresses five core challenges in dynamic data environments—data drift, concept drift, catastrophic forgetting, skewed learning, and network adaptability. Method: It systematically surveys over 120 state-of-the-art evolutionary machine learning (EML) works, integrating online learning, incremental learning, continual learning, meta-learning, dynamic pruning, and ensemble distillation to establish a multi-paradigm evaluation framework covering supervised, unsupervised, and semi-supervised settings. Contribution/Results: The work introduces the first unified analytical framework for EML, clarifies challenge taxonomies, uncovers synergistic mechanisms among adaptive neural architectures, meta-learning, and ensemble strategies, and identifies critical gaps in robustness, ethics, and scalability. It delivers a comprehensive EML methodology landscape, a curated collection of mainstream benchmarks and evaluation metrics, and system design principles tailored for industrial deployment—providing both theoretical foundations and practical guidance for building dynamic AI systems.
This work addresses the challenge of continual learning in dynamic environments characterized by non-stationary data streams by proposing Streamed Continual Learning (SCL), a novel framework that unifies the paradigms of continual learning and streaming machine learning for the first time. SCL integrates knowledge retention mechanisms from continual learning with online updating strategies from streaming learning to construct an efficient and adaptive unified architecture. The study delineates the core characteristics and technical pathways of SCL, thereby fostering synergistic innovation between the two fields and laying a theoretical foundation for developing general-purpose adaptive intelligent systems capable of operating effectively in dynamic environments.
This work addresses the challenge faced by agents in real-world scenarios of simultaneously adapting rapidly to concept drift while avoiding catastrophic forgetting in continual learning. To this end, we formally introduce and define the Streaming Continual Learning (SCL) paradigm, which integrates knowledge retention mechanisms from continual learning with the online updating and concept drift detection capabilities of streaming machine learning. This framework bridges two previously distinct research communities and fosters the development of hybrid approaches that balance rapid adaptation with long-term memory. Empirical results demonstrate that conventional methods from either continual learning or streaming machine learning alone struggle to achieve this balance, whereas SCL significantly outperforms them in non-stationary data streams, effectively reconciling fast adaptation with stable knowledge retention.
To address the high computational cost of SVD approximation updates in streaming data and the scalability limitations of existing incremental/truncated SVD methods at large truncation ranks, this paper proposes a low-rank update framework based on Bidirectional Diagonal Decomposition (BDD). Our method enables efficient rank-$r$ updates via three key innovations: (1) a compact Householder transformation reducing memory usage by 50%; (2) Givens rotations enabling $O(r^2)$-complexity rank-$r$ updates; and (3) a hybrid sparse-plus-low-rank separation strategy for accurate and scalable matrix approximation. Experiments on recommendation systems and network subspace tracking demonstrate that our approach significantly outperforms LAPACK SVD and state-of-the-art incremental SVD methods—achieving superior accuracy, real-time performance even at high truncation ranks, and balanced throughput–precision trade-offs.
This work addresses the challenge of rapid dynamic shifts in streaming time series caused by abrupt environmental changes or varying input delays. The authors propose a system tensor representation based on Markov parameter sequences, modeling the streaming data as a dynamic mixture of delay systems. By constructing fixed-length tensor summaries that jointly encode system dynamics and input–output delay characteristics, the method enables efficient compression and retrieval of historical patterns through tensor decomposition. Within an online learning framework, the system dynamically selects the optimal submodel to match the current state, achieving strong adaptability to nonstationary time series while maintaining low memory overhead. Experimental results on real-world datasets demonstrate that the proposed approach significantly outperforms existing methods in both prediction accuracy and adaptation speed, particularly under highly nonstationary conditions.
This work addresses the challenge of simultaneously achieving real-time responsiveness, stability, and tracking performance in nonstationary data streams, where conventional adaptive algorithms often fall short. Focusing on the Momentum Least Mean Squares (MLMS) algorithm, the study establishes, for the first time, rigorous stability conditions, tracking error bounds, and dynamic regret bounds under time-varying stochastic linear systems. The analysis overcomes the significant theoretical difficulty posed by second-order products of stochastic matrices induced by the momentum term. By integrating tools from stochastic vector difference equations, nonstationary time series modeling, and online learning theory, the paper demonstrates that MLMS exhibits both rapid adaptability and robust tracking capability. Empirical evaluations on synthetic and real-world data streams confirm that MLMS significantly outperforms the classical LMS algorithm.
Addressing the coupled challenges of concept drift and catastrophic forgetting in unbounded data streams, this paper proposes a unified Stream-based Continual Learning (SCL) framework. Methodologically, it introduces a Large-scale Context-Aware Tabular Model (LTM) as a central hub—first rigorously demonstrating that such a model simultaneously satisfies streaming constraints (e.g., bounded memory, low-latency inference) and continual learning requirements (e.g., experience replay). The approach is grounded in two principled objectives: distribution matching and distribution compression, jointly optimizing plasticity (adaptation to novel distributions), stability (preservation of prior knowledge), memory diversity, and retrieval priority. It integrates online stream summarization, dynamic distribution matching, diversity-driven memory compression, and adaptive retrieval. Experiments across multiple SCL benchmarks show significant forgetting mitigation, sub-millisecond inference latency, strict memory boundedness, and superior overall performance over state-of-the-art continual learning and streaming learning methods.
This work addresses the scalability and adaptability limitations of traditional Operator Inference methods, which require full data loading and lack support for online updates—challenges that hinder their application to large-scale systems under memory constraints. To overcome these issues, we introduce, for the first time, a streaming learning framework into Operator Inference, proposing a non-intrusive model order reduction approach that leverages incremental singular value decomposition (Incremental SVD) and recursive least squares (RLS). This method enables the online construction and adaptive updating of low-dimensional dynamical models directly from continuous data streams. Experimental results demonstrate that the proposed technique achieves accuracy comparable to batch-based methods while reducing memory consumption by over 99%, achieving a dimensionality compression ratio exceeding 31,000×, and accelerating predictions by several orders of magnitude.
To address catastrophic forgetting in memory-constrained streaming learning, this paper proposes a unified continual learning framework based on state replay, applicable to generative (autoencoding), time-series forecasting, and classification tasks. Unlike naive sequential fine-tuning or black-box replay, we formulate state replay as a joint optimization objective, enabling cooperative parameter updates over new and replayed samples via stochastic gradient methods; theoretical analysis grounded in gradient alignment reveals necessary conditions for effective forgetting mitigation. Evaluated across six heterogeneous and stationary streaming settings—constructed from Rotated MNIST, Electricity, and Airlines datasets—the method reduces average forgetting by 2–3× under heterogeneous multi-task streams, while matching fine-tuning performance on stationary streams. Our key contributions are: (i) the first theoretical analysis framework for state replay that unifies generative and discriminative tasks, and (ii) empirical validation of its effectiveness and robustness as a strong baseline for streaming continual learning.