PaLoRA: Paced Low-Rank Adaptation for Continual Learning

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of theoretical justification for fixed small learning rates in LoRA-based continual learning and the exacerbation of catastrophic forgetting as tasks accumulate. We reveal a knowledge leakage mechanism under finite precision and derive an optimal gradient scaling law based on effective rank. Accordingly, we propose an adaptive pacing rule that dynamically adjusts gradient magnitudes via an anisotropic leakage model. By integrating adaptive SVD truncation for historical knowledge compression, null-space projection to constrain update directions, and rank-aware step-size control, our method achieves an optimal stability-plasticity balance. Evaluated on a 50-task benchmark including ImageNet-A/R, it yields a 4% accuracy improvement over existing approaches, demonstrating particular strength in periodic scenarios.
📝 Abstract
LoRA-based continual learning methods mitigate catastrophic forgetting through various mechanisms, yet nearly all complement these with small learning rates as a heuristic to restrict gradient scaling magnitude. Such fixed heuristics lack theoretical guidance on how the strength of this restriction should evolve as tasks accumulate. We reveal that even under directional constraints such as nullspace projection, finite-precision updates inevitably leak into the subspace of accumulated prior knowledge along multiple directions. While small learning rates attenuate such leakage, they cannot prevent the accumulated forgetting from intensifying as the effective rank of historical knowledge grows. We show that the optimal magnitude restriction should adaptively increase with this effective rank to balance stability and plasticity, i.e., preservation of previous knowledge and acquisition of new task information. Under an anisotropic leakage model, we derive a pacing law $s^*=\sqrt{R/c}$ that characterizes the optimal scaling of gradient steps, i.e., the magnitude restriction itself, where $R$ is the effective rank of past updates. Based on this insight, we propose PaLoRA, which compresses historical knowledge via adaptive SVD truncation, projects gradients onto the nullspace of prior tasks, and applies rank-aware adaptive pacing. Experiments demonstrate consistent improvements over prior methods, with particularly strong performance in long-horizon settings, achieving substantial gains of 4% accuracy on challenging 50-task ImageNet-A and ImageNet-R benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Continual Learning
Catastrophic Forgetting
Low-Rank Adaptation
Stability-Plasticity Dilemma
Effective Rank
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continual Learning
Low-Rank Adaptation
Adaptive Pacing
Nullspace Projection
Catastrophic Forgetting
🔎 Similar Papers
No similar papers found.