continual learning

Designs, implements, and analyzes methods and training procedures for updating models on streams or sequences of tasks or datasets while preventing catastrophic forgetting, including parameter-efficient and few-parameter adaptation, sequential fine-tuning, replay/rehearsal storage and replay strategies, regularization and aggregation techniques, and incremental learning schedules. Builds benchmarks and evaluation protocols to measure forgetting and sequential task retention, long-sequence robustness, and overall continual-learning performance.

continuallearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.73
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Mitigating Catastrophic Forgetting in Streaming Generative and Predictive Learning via Stateful Replay

Nov 22, 2025
WD
Wenzhang Du
🏛️ Mahanakorn University of Technology

To address catastrophic forgetting in memory-constrained streaming learning, this paper proposes a unified continual learning framework based on state replay, applicable to generative (autoencoding), time-series forecasting, and classification tasks. Unlike naive sequential fine-tuning or black-box replay, we formulate state replay as a joint optimization objective, enabling cooperative parameter updates over new and replayed samples via stochastic gradient methods; theoretical analysis grounded in gradient alignment reveals necessary conditions for effective forgetting mitigation. Evaluated across six heterogeneous and stationary streaming settings—constructed from Rotated MNIST, Electricity, and Airlines datasets—the method reduces average forgetting by 2–3× under heterogeneous multi-task streams, while matching fine-tuning performance on stationary streams. Our key contributions are: (i) the first theoretical analysis framework for state replay that unifies generative and discriminative tasks, and (ii) empirical validation of its effectiveness and robustness as a strong baseline for streaming continual learning.

Evaluating stateful replay on heterogeneous multitask and time-based data streamsMitigating catastrophic forgetting in streaming learning under memory constraintsStudying replay mechanisms across generative and predictive learning objectives

Replay Can Provably Increase Forgetting

Jun 04, 2025
YM
Yasaman Mahdaviyeh
🏛️ Columbia University | NVIDIA | New York University | Stanford University

Sample replay is widely employed in continual learning to mitigate catastrophic forgetting of old tasks, yet its efficacy remains poorly understood. Method: We theoretically analyze replay in an idealized, noiseless, overparameterized linear regression setting and conduct empirical validation via SGD-trained neural networks on standard benchmarks. Contribution/Results: We establish, for the first time, that replay can *worsen* both worst-case and expected forgetting—demonstrating non-monotonic and even detrimental effects. The core mechanism is a coupling between task subspace geometry and replay sample selection: when replayed samples deviate from the principal directions of old-task subspaces, they amplify parameter drift and induce negative transfer. Our theoretical forgetting bounds and extensive experiments consistently reproduce and confirm this phenomenon. This work challenges the assumption that replay is universally beneficial, revealing its effectiveness to be critically contingent on task geometry and replay strategy—providing essential theoretical guidance and caution for designing robust replay mechanisms in continual learning.

Analyzes sample replay's impact on forgetting in continual learningDemonstrates harmful replay scenarios in linear and neural modelsIdentifies non-monotonic forgetting despite sufficient replay samples

This work addresses the challenge of catastrophic forgetting in large language models during continual fine-tuning, a problem inadequately mitigated by existing replay methods that either rely on heuristic rules or incur high computational costs. Inspired by human memory mechanisms, the authors propose the Memory Strength–aware Sample Replay (MSSR) framework, which dynamically optimizes both the content and timing of replayed samples through sample-level memory strength estimation and adaptive replay scheduling. By integrating experience replay, memory strength modeling, and a lightweight scheduling algorithm, MSSR effectively balances the mitigation of forgetting with rapid adaptation to new tasks. Extensive experiments across three backbone models and eleven sequential tasks demonstrate that MSSR significantly outperforms current replay strategies, with particularly notable gains on reasoning-intensive and multiple-choice tasks.

catastrophic forgettingcontinual fine-tuningexperience replay

To mitigate catastrophic forgetting in multi-stage fine-tuning—exacerbated by computational constraints on foundation models—this paper proposes a lightweight replay sample selection paradigm requiring no additional forward passes. Our method introduces the mix-cd sampling mechanism, which estimates the density distribution of “collateral damage” samples via prediction consistency analysis, enabling precise identification and efficient filtering of highly forgettable instances. Unlike conventional approaches, it eliminates reliance on fixed-size memory buffers or auxiliary model inference, achieving superior knowledge retention under strict computational budgets. Experiments demonstrate that our approach attains state-of-the-art continual learning performance with significantly lower overhead across multiple benchmarks. The implementation is publicly available.

Efficiently estimate sample density without additional inferencesMitigate catastrophic forgetting in multi-stage fine-tuningPrioritize rehearsal of collateral damage samples

Catastrophic forgetting remains a fundamental challenge in continual learning. This work investigates sample-level forgetting sensitivity and identifies a strong correlation between learning order and forgetting severity: samples learned earlier exhibit greater resistance to forgetting. Motivated by this finding, we propose the “Goldilocks” sampling principle—selecting only moderately learned samples for rehearsal while excluding those learned too quickly or too slowly. We further design a training-dynamics-based model to estimate per-sample learning speed and integrate it into a dynamic buffer update mechanism. Our approach is lightweight and seamlessly compatible with mainstream rehearsal methods (e.g., ER, A-GEM). Extensive experiments on Split-CIFAR10/100 and Split-ImageNet demonstrate state-of-the-art performance: average accuracy improves by 2.1–3.7%, and forgetting rates decrease significantly. To our knowledge, this is the first work to quantitatively establish the relationship between learning timing and forgetting, introducing a novel sample-aware paradigm for continual learning.

Investigating replay buffer composition impact on forgettingPredicting susceptibility to catastrophic forgetting in neural networksProposing Speed-Based Sampling to improve continual learning performance

Latest Papers

What's happening recently
View more

This work addresses catastrophic forgetting in continual learning under non-stationary data streams by proposing the COLD framework, which introduces, for the first time, the Drift-Plus-Penalty stochastic optimization method from control theory into this domain. COLD formulates forgetting as a controlled dynamic process, employing virtual queues to track performance deviations on historical tasks and jointly minimizing the current task loss and queue drift at each optimization step. This mechanism explicitly governs the stability-plasticity trade-off. The framework provides theoretical guarantees on stability and convergence, and achieves significantly superior performance over state-of-the-art methods on standard benchmarks, enabling controllable and efficient suppression of catastrophic forgetting.

catastrophic forgettingcontinual learningnonstationary data streams

This work addresses the challenge in continual learning where training on new tasks often leads to significant performance degradation on previously learned tasks, a phenomenon known as catastrophic forgetting. Existing approaches struggle to identify the directions in output space most susceptible to interference. Building upon the Neural Tangent Kernel (NTK) framework, this study derives a closed-form expression for prediction drift on old tasks in function space, establishing—for the first time—an exact analytical link between forgetting vectors and cross-task kernels. The analysis reveals that forgetting exhibits a low-rank structure and introduces a Kronecker scaling law for the "fragile rank," clarifying its relationship with NTK overlap theory. Under a PEFT-CL setting with frozen backbones and linear heads, the derived expression predicts forgetting with numerical precision, enabling a spectrally regularized method that effectively suppresses interference in output space.

Catastrophic ForgettingContinual AdaptationFunction-Space

This work addresses a critical gap in continual learning research: while most existing methods focus on mitigating catastrophic forgetting, they largely overlook the conditions under which forward transfer—where knowledge from past tasks benefits new ones—can be effectively realized. The paper introduces, for the first time, a systematic three-condition framework that characterizes when forward transfer is feasible and proposes Transfer-Selective Replay (TSR), a novel method that leverages a zero-training-overhead task signature mechanism to automatically identify and replay only those historical samples beneficial to the current task. TSR integrates knowledge distillation to preserve performance on previous tasks while explicitly promoting forward transfer as a first-class objective. Experiments demonstrate that TSR significantly enhances forward transfer across both homogeneous and heterogeneous task sequences and consistently outperforms existing replay-based baselines, especially under limited replay budgets.

catastrophic forgettingcontinual learningforward transfer

This study investigates the impact of task granularity ordering on catastrophic forgetting in continual learning, presenting the first systematic evaluation of three learning strategies—coarse-to-fine, fine-to-coarse, and flat learning—on CIFAR-100. Leveraging Elastic Weight Consolidation (EWC), the authors assess model performance using accuracy, F1 score, and continual learning–specific metrics to quantify the retention of previously acquired knowledge. The findings demonstrate that initializing learning with coarse-grained categories before introducing fine-grained tasks significantly mitigates catastrophic forgetting and enhances backward transfer. This suggests that incorporating hierarchical priors to construct stable representations offers an effective principle for designing task sequences in incremental learning scenarios, thereby providing a novel strategy for optimizing continual learning systems.

catastrophic forgettingcontinual learningknowledge retention

This work addresses the challenges of catastrophic forgetting and task-specific knowledge dilution in continual fine-tuning of large language models. Existing approaches typically rely on experience replay or task-specific adapters, incurring substantial computational and storage overhead. To overcome these limitations, the authors propose a novel paradigm that requires neither replay nor additional adapter modules. Their method employs a brief warm-up fine-tuning phase, followed by identification of a core subset of parameters per task using parameter importance metrics—such as L2 norm and Fisher information—and task-specificity analysis based on cosine similarity of update directions. During subsequent training, only this critical parameter subset is updated while the rest remain frozen to preserve prior knowledge. Extensive experiments demonstrate that this approach significantly outperforms current state-of-the-art methods across multiple benchmarks, confirming its effectiveness for large-scale models under resource constraints and its transferability across different model sizes.

catastrophic forgettingcontinual fine-tuninglarge language models

Hot Scholars

DW

Da-Wei Zhou

Associate Researcher, Nanjing University
Incremental LearningContinual LearningOpen-World LearningModel Reuse
DD

Davide Dalle Pezze

University of Padua
Deep LearningIndustry 4.0Continual LearningVisual Anomaly Detection
DG

Dong Gong

University of New South Wales (UNSW)
Computer VisionImage ProcessingMachine Learning
HZ

Huiping Zhuang

Associate Professor, South China University of Technology
Continual LearningMulti-ModalEmbodied AILarge Model
HJ

Han-Jia Ye

Nanjing University
Machine LearningData MiningMetric LearningMeta-Learning