Score
Designs and evaluates continual learning training procedures that enforce constraints on what the model relies on so new task updates preserve earlier checkpoints’ predictions and evidence patterns. Implementations typically freeze a previous checkpoint as a reference and optimize the new-task loss with added reliance constraints to maintain prior behavior while avoiding added inference-time cost.
This work addresses catastrophic forgetting in continual learning under non-stationary data streams by proposing the COLD framework, which introduces, for the first time, the Drift-Plus-Penalty stochastic optimization method from control theory into this domain. COLD formulates forgetting as a controlled dynamic process, employing virtual queues to track performance deviations on historical tasks and jointly minimizing the current task loss and queue drift at each optimization step. This mechanism explicitly governs the stability-plasticity trade-off. The framework provides theoretical guarantees on stability and convergence, and achieves significantly superior performance over state-of-the-art methods on standard benchmarks, enabling controllable and efficient suppression of catastrophic forgetting.
This work addresses catastrophic forgetting in continual learning caused by distribution shifts across tasks, particularly when tasks exhibit dependencies. The authors posit that data from the current task can be modeled as a nonlinear transformation of data from previous tasks, thereby formalizing a task-dependency structure. Building on this assumption, they integrate techniques from nonlinear regression, experience replay, and knowledge distillation to derive, for the first time, a non-vacuous estimation error bound with practical significance. This theoretical framework provides the first rigorous statistical recovery guarantee for continual learning methods that incorporate memory replay and multiple regularization strategies, substantially enhancing the interpretability and reliability of such algorithms.
This work addresses the limitation of traditional continual learning, which overly emphasizes preserving old knowledge to approximate joint training while neglecting real-time adaptation to new tasks in non-stationary environments. The problem is reformulated as an online optimization framework, with average lifelong error as the performance metric, and a notion of transfer efficiency is introduced to characterize the trade-off between stability and transient error induced by historical knowledge. Drawing on critical task duration theory, the study identifies conditions under which past knowledge shifts from beneficial to detrimental. Building on this insight, the paper proposes a predictive continual learning paradigm that dynamically models future tasks to optimize expected performance. Integrating online learning, transfer efficiency analysis, and convergence theory, an algorithm based on task-window interpolation is developed and validated on image classification and reinforcement learning benchmarks, demonstrating significant superiority over both joint training and independent learning, especially under distribution shift.
A core challenge in continual learning is enabling AI agents to efficiently accumulate knowledge and progressively enhance skills over extended operational lifetimes. Method: This work introduces the first formal continual learning definition that explicitly incorporates computational resource constraints, modeling it as a constrained reinforcement learning process to unify objectives, constraints, and evaluation. We propose an analytically tractable RL-based framework integrating computational complexity analysis, information-theoretic principles, and dynamic resource modeling. Contribution/Results: We establish the first theoretical framework jointly optimizing computational efficiency and knowledge accumulation capacity. It provides falsifiable hypotheses and benchmarking tools for rigorous empirical validation. Our approach significantly advances the mathematical rigor, scalability, and real-world deployability of continual learning systems—enabling principled trade-offs between learning performance, memory footprint, and inference latency under bounded resources.
Current continual learning (CL) evaluation protocols suffer from critical flaws—hyperparameter tuning and evaluation are conducted within the same scenario, leading to systematic overestimation of CL capability and employing unrealistic, non-deployable tuning practices. Method: We propose the Generalized Two-stage Evaluation Protocol (GTEP), which strictly decouples hyperparameter optimization (performed solely on a source dataset) from performance evaluation (conducted on a target dataset), thereby enforcing cross-dataset generalization under structurally identical tasks. Contribution/Results: Extensive experiments—over 8,000 runs across CIFAR and ImageNet variants—under both pre-trained and non-pre-trained settings within a class-incremental learning framework demonstrate that mainstream SOTA methods suffer 30–50% average performance degradation under GTEP. This reveals their lack of robustness across deployment scenarios and establishes GTEP as a more rigorous, realistic benchmark for trustworthy continual learning.
This work addresses the challenge that AI agents with frozen weights after deployment struggle to learn continuously from experience, often failing on repeated tasks. The authors propose a continual learning mechanism leveraging external memory, which distills minimal feedback—either a single-bit outcome or natural language corrections—from each interaction into retrievable rules. Integrated with retrieval-augmented generation (RAG) and frozen large language models (e.g., Mistral Large, Claude Sonnet 5), this approach enables performance improvement without fine-tuning. The method demonstrates, for the first time, that extremely sparse feedback alone can drive sustained enhancement in frozen models and supports memory transfer across models. On the τ-bench banking tasks, it achieves success rates 1.6× (outcome-only feedback) and 2.6× (with corrections) higher than baseline, resolving 22 out of 84 tasks on which the baseline completely fails.
Traditional continual learning is constrained by a parameter-centric paradigm, limiting its capacity to meet system-level adaptation demands in dynamic environments. This work proposes a “Tri-Axis Framework” (When, How, Where), offering a unified perspective that reorients continual learning beyond mere parameter updates toward external architectures and inference-time adaptation. By integrating off-policy/on-policy learning, test-time training, external memory systems, and skill repositories, the framework transcends the limitations of static parameter spaces and gradient-based optimization. A systematic review elucidates the field’s evolutionary trajectory and highlights pivotal challenges and future directions inherent in this paradigm shift.
This study addresses the dual challenges of emerging domains and data drift faced by large language models in dynamic environments. The authors decompose continual learning into spatial (new domains) and temporal (data drift) dimensions and propose the first mechanism-agnostic, unified evaluation protocol. Within this consistent framework, they systematically compare the adaptability of diverse approaches—including prompt engineering (e.g., GEPA, ACE), supervised fine-tuning (SFT, SDFT), online reinforcement learning (GRPO, SDPO), and context compression (Cartridges, In-place TTT). Their analysis reveals that effective adaptation depends on aligning update mechanisms with specific environmental dynamics: online reinforcement learning excels at knowledge updating yet is sensitive to noise; distillation-based methods offer stability but struggle to correct outdated facts; and prompt-based strategies respond rapidly but exhibit limited generalization.