🤖 AI Summary
This work addresses catastrophic forgetting in large language models under continual learning, where existing LoRA-based approaches often rely on task identifiers or trainable routing modules, compromising parameter efficiency and hindering task-agnostic inference. The authors propose a gradient-free task selection mechanism that eliminates the need for trainable routers by leveraging the distributional separation of pooled token embeddings from frozen embedding layers, enabling automatic adapter selection via a gradient-free Gaussian Mixture Model. Additionally, LoRA parameters are constrained to the principal subspace of the pre-trained weight’s SVD decomposition and regularized with orthogonality constraints to minimize interference across tasks. Evaluated across five model scales and two continual learning benchmarks, the method achieves near-zero forgetting without rehearsal, significantly reduces per-task parameters, and sets a new state-of-the-art performance.
📝 Abstract
Large language models generalize well to individual tasks but lack an inherent mechanism for learning them sequentially, leading to catastrophic forgetting. To mitigate this, LoRA-based continual learning methods allocate a separate low-rank adapter per task, yet existing approaches either require task identity at inference or sum all adapters indiscriminately, letting irrelevant branches distort the output. Recent gating-based solutions route inputs to the correct adapter but introduce trainable parameters that themselves need protection against forgetting. In this work, we observe that pooled token embeddings from a frozen LLM embedding layer already separate task distributions throughout the learning sequence. A Gaussian mixture model fitted on these embeddings, without any gradient-based training, is sufficient for task-agnostic adapter selection at test time. This eliminates the need for a learned gating module. On the adapter side, constraining each task's parameters to the principal subspace of the pretrained weights via SVD yields a compact latent-space parameterization. Within this subspace, orthogonal regularization directly controls inter-task interference. The resulting system, Latent-LoRA, is replay-free, requires no trainable routing component, and uses substantially fewer parameters per task. Experiments across five model scales and two established continual learning benchmarks show state-of-the-art performance with near-zero forgetting.