Memory-Statistics Tradeoff in Continual Learning with Structural Regularization

📅 2025-04-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the fundamental trade-off between memory overhead and statistical efficiency in continual learning, focusing on the two-task linear regression setting under random design. To mitigate catastrophic forgetting, we propose a generalized ℓ₂-structured regularization method that leverages the Hessian structure of prior tasks. Our theoretical analysis provides the first rigorous characterization of the quantitative trade-off between memory complexity—measured by the dimension of stored vectors—and excess risk. We prove that unregularized continual learning inevitably leads to statistical collapse, whereas our method achieves excess risk convergence at the optimal rate of joint training, attaining statistical efficiency comparable to full-data joint estimation, using only O(d) memory (where d is the feature dimension). This work establishes the first regularization framework for continual learning that is simultaneously structure-aware, theoretically tight, and computationally feasible.

Technology Category

Machine Learning: Life-Long and Continual LearningNatural Language Processing: Learning & Optimization for NLPSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
We study the statistical performance of a continual learning problem with two linear regression tasks in a well-specified random design setting. We consider a structural regularization algorithm that incorporates a generalized $ell_2$-regularization tailored to the Hessian of the previous task for mitigating catastrophic forgetting. We establish upper and lower bounds on the joint excess risk for this algorithm. Our analysis reveals a fundamental trade-off between memory complexity and statistical efficiency, where memory complexity is measured by the number of vectors needed to define the structural regularization. Specifically, increasing the number of vectors in structural regularization leads to a worse memory complexity but an improved excess risk, and vice versa. Furthermore, our theory suggests that naive continual learning without regularization suffers from catastrophic forgetting, while structural regularization mitigates this issue. Notably, structural regularization achieves comparable performance to joint training with access to both tasks simultaneously. These results highlight the critical role of curvature-aware regularization for continual learning.
Problem

Research questions and friction points this paper is trying to address.

Trade-off between memory complexity and statistical efficiency in continual learning
Mitigating catastrophic forgetting via curvature-aware structural regularization
Comparing performance of structural regularization versus naive continual learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structural regularization with Hessian-tailored ell_2
Trade-off between memory complexity and risk
Curvature-aware regularization prevents catastrophic forgetting
🔎 Similar Papers
No similar papers found.