🤖 AI Summary
This study addresses the unclear mechanisms underlying catastrophic forgetting induced by LoRA in continual learning. Leveraging high-dimensional online learning theory and teacher-student solvable models, we derive asymptotically exact dynamical equations for LoRA. The analysis quantifies the trade-off between interference suppression and adaptation speed in low-rank updates, elucidating the geometric role of adapter rank and the theoretical advantages of state-dependent masking strategies. Our results demonstrate that the proposed approach significantly mitigates forgetting while preserving plasticity for new tasks. Theoretical predictions exhibit strong agreement with both numerical simulations and MNIST experiments, establishing a rigorous dynamical foundation for parameter-efficient continual learning.
📝 Abstract
Despite the widespread use of Low-Rank Adaptation (LoRA), little is known about its dynamics in continual learning and the mechanisms by which low-rank updates affect catastrophic forgetting. We provide an asymptotically exact dynamical characterization of LoRA in a solvable two-task teacher-student model. In the high-dimensional online-learning limit, we derive a closed system of ordinary differential equations for a finite set of macroscopic order parameters, yielding exact expressions for the generalization errors throughout both the initial Task 1 learning phase and the subsequent LoRA fine-tuning on Task 2. The theory quantitatively matches finite-dimensional simulations and exposes two characteristic effects of LoRA: low-rank adaptation reduces interference with features learned on the first task, but its initialization slows adaptation to the second task. Building on this mechanistic picture, we analyze a state-dependent masking strategy that freezes hidden units carrying the strongest first-task representations and restricts adaptation to the complementary subspace. This structural partitioning markedly reduces forgetting, while preserving plasticity on the new task. Our framework further clarifies the role of adapter rank: transfer improves only up to the intrinsic dimensionality of the target task and saturates beyond it, while forgetting continues to grow with rank. These results provide a dynamical and geometric account of how low-rank adaptation organizes information across sequential tasks and are qualitatively reproduced on a sequential MNIST benchmark.