Score
Designs, implements, and evaluates regularization-based continual learning methods that mitigate catastrophic forgetting by penalizing updates to model parameters deemed important for previous tasks. This includes algorithms such as elastic weight consolidation that estimate parameter importance (e.g., via Fisher information or similar approximations) and add quadratic or other penalty terms to the loss to constrain later-task training.
This work addresses catastrophic forgetting in continual learning by proposing and systematically evaluating an improved implementation of Elastic Weight Consolidation (EWC). We reproduce EWC on the PermutedMNIST and RotatedMNIST benchmarks, augmenting the analysis with dropout integration and comprehensive hyperparameter sensitivity studies. Results demonstrate that EWC effectively balances knowledge retention and new-task adaptation: it substantially mitigates forgetting—outperforming unregularized SGD baselines—and enhances overall continual learning stability and generalization, albeit with a modest reduction in convergence speed on new tasks. Our study confirms EWC’s robustness across canonical non-stationary task sequences and reveals synergistic effects between regularization strength and architectural choices (e.g., dropout rate) on continual learning efficacy. Crucially, this work provides reproducible empirical evidence supporting lightweight, regularization-based continual learning approaches.
This work addresses a fundamental flaw in Elastic Weight Consolidation (EWC) and its variants for continual learning, where biased importance estimates—stemming from vanishing gradients and redundant protection in the Fisher Information Matrix (FIM)—exacerbate catastrophic forgetting. The study is the first to reveal this core deficiency in weight importance evaluation and introduces a Logits Reversal operation that effectively rectifies the FIM computation without requiring complex architectural modifications. This simple yet effective adjustment yields more accurate parameter importance estimates. Extensive experiments across diverse continual learning benchmarks demonstrate that the proposed method consistently outperforms original EWC and its extensions, achieving superior model performance while significantly mitigating catastrophic forgetting.
This work addresses the dual challenges of inter-task interference and computational overhead in parameter-efficient continual learning (PECL). The authors propose integrating Elastic Weight Consolidation (EWC) regularization into the Low-Rank Adaptation (LoRA) framework by imposing importance-based constraints on shared low-rank updates. This approach effectively balances model stability and plasticity without increasing parameter count or inference cost. The key innovation lies in leveraging low-rank representations to efficiently approximate full-dimensional parameter importance, enabling lightweight weight regularization. Experimental results demonstrate that the method significantly outperforms existing low-rank approaches across multiple continual learning benchmarks, achieving a superior trade-off between stability and plasticity.
This work investigates the fundamental trade-off between memory overhead and statistical efficiency in continual learning, focusing on the two-task linear regression setting under random design. To mitigate catastrophic forgetting, we propose a generalized ℓ₂-structured regularization method that leverages the Hessian structure of prior tasks. Our theoretical analysis provides the first rigorous characterization of the quantitative trade-off between memory complexity—measured by the dimension of stored vectors—and excess risk. We prove that unregularized continual learning inevitably leads to statistical collapse, whereas our method achieves excess risk convergence at the optimal rate of joint training, attaining statistical efficiency comparable to full-data joint estimation, using only O(d) memory (where d is the feature dimension). This work establishes the first regularization framework for continual learning that is simultaneously structure-aware, theoretically tight, and computationally feasible.
To address the coupled generalization-forgetting error problem arising from catastrophic forgetting in continual learning, this paper pioneers modeling the temporal evolution of the loss function as a parabolic partial differential equation (PDE), with the memory buffer serving as a dynamic boundary condition—explicitly capturing long-range dependencies and error propagation. Leveraging the intrinsic physical regularity of PDEs, we formulate a spatiotemporal constrained optimization framework driven by boundary conditions, enabling analyzable and interpretable dynamic regularization. Theoretically, we derive a tight coupled bound on forgetting and generalization errors. Empirically, our method significantly reduces forgetting across multiple standard benchmarks, and the theoretical error bound closely aligns with observed performance—demonstrating the effectiveness, stability, and analytical tractability of PDE-based regularization.
This work addresses catastrophic forgetting in continual learning caused by distribution shifts across tasks, particularly when tasks exhibit dependencies. The authors posit that data from the current task can be modeled as a nonlinear transformation of data from previous tasks, thereby formalizing a task-dependency structure. Building on this assumption, they integrate techniques from nonlinear regression, experience replay, and knowledge distillation to derive, for the first time, a non-vacuous estimation error bound with practical significance. This theoretical framework provides the first rigorous statistical recovery guarantee for continual learning methods that incorporate memory replay and multiple regularization strategies, substantially enhancing the interpretability and reliability of such algorithms.
This work addresses the degradation of generalization performance in high-dimensional continual multi-task linear regression caused by label noise and overparameterization. For isotropic L2-regularized linear models, the authors derive a closed-form expression for the expected generalization loss under an arbitrary linear teacher model and establish, for the first time, that the optimal fixed regularization strength grows with the number of tasks \(T\) at a rate of \(T / \ln T\). This result uncovers a novel mechanism by which L2 regularization enhances generalization—by suppressing the adverse effects of label noise. The theoretical analysis leverages tools from high-dimensional statistics and is corroborated by experiments on both linear models and neural networks, demonstrating that the proposed regularization strategy significantly improves generalization in continual learning settings and offers practical design principles for real-world systems.
This work addresses the fundamental challenge in continual learning of balancing the acquisition of new knowledge with the retention of previously learned information under limited model capacity, where controlled forgetting plays a pivotal role. The authors propose FADE, a novel method that, for the first time, integrates online approximate meta-gradient descent into a weight decay mechanism to dynamically and fine-grainedly adjust the decay rate of individual parameters. Built upon meta-gradient derivation in an online linear setting, FADE focuses on the final layer of neural networks and jointly optimizes decay rates with adaptive step sizes. Empirical results on online tracking and streaming classification tasks demonstrate that FADE automatically discerns the stability requirements of different parameters and significantly outperforms fixed decay strategies.
This work addresses the lack of theoretical understanding in existing continual learning methods regarding how the magnitude of parameter updates influences forgetting and generalization. Formalizing forgetting as knowledge degradation caused by task-specific parameter drift, the study introduces an optimization framework that constrains parameter updates. It innovatively unifies frozen training and initialization-based training within a single theoretical framework and proposes a hybrid strategy that adaptively adjusts update magnitudes based on gradient directions. Through parameter space analysis and constrained optimization, the method significantly reduces catastrophic forgetting while enhancing generalization performance in deep neural networks, outperforming standard training approaches.
In continual learning, models suffer from catastrophic forgetting when acquiring new tasks. This paper introduces task sequence ordering as an independent intervention dimension—first proposed in the literature—and presents a zero-shot scoring method to identify optimal task orderings. By modeling task similarity and adapting neural architecture search (NAS) heuristics, our approach estimates cross-task transfer potential without fine-tuning. It seamlessly integrates with mainstream continual learning strategies—including Elastic Weight Consolidation (EWC) and experience replay—yielding substantial mitigation of forgetting: average accuracy improves by 3.2–7.8% across multiple benchmarks, while retention of previously learned tasks increases by 12.4–29.6%. Moreover, it enhances the robustness and generalization of diverse baseline methods. The core innovation lies in decoupling task-order optimization from model parameter updates, establishing a novel paradigm for continual learning that separates sequencing logic from learning dynamics.