Score
Design, implement, and evaluate methods that apply LoRA-style low-rank adapter modules to enable sequential (continual) fine-tuning of pretrained models with the backbone frozen, including adapter parameterization, update rules, storage and inference-time composition. Build and analyze strategies to prevent catastrophic forgetting and retain prior-task performance across adapter updates—e.g., consolidation, replay, regularization, adapter selection, and evaluation protocols for continual learning with LoRA.
This work addresses the fragmented landscape of LoRA variants, which currently lack a unified taxonomy, theoretical framework, and standardized implementation and evaluation protocols. To this end, we propose the first four-dimensional classification scheme grounded in rank structure, optimization dynamics, initialization strategies, and MoE integration, offering a cohesive theoretical perspective. We further develop LoRAFactory, a modular codebase enabling systematic experimentation across diverse tasks—including natural language generation, natural language understanding, and image classification—through large-scale empirical studies. Our findings reveal that the original LoRA, when equipped with well-tuned hyperparameters, matches or surpasses most existing variants, while exhibiting pronounced sensitivity to learning rate choices. These results underscore LoRA’s robustness and efficacy, establishing it as a standardized benchmark for parameter-efficient fine-tuning.
To address the high computational cost and poor generalization in efficient adaptation of foundation models, this paper presents the first systematic survey of Low-Rank Adaptation (LoRA) extensions across broad classes of foundation models—including multimodal and scientific computing models. We propose a unified taxonomy that integrates matrix low-rank decomposition, modular adapter design, gradient-constrained optimization, and cross-task transfer analysis—thereby identifying key theoretical gaps and charting a new direction toward robustness-aware modeling. Covering over 100 state-of-the-art works, we uncover common mechanisms underlying LoRA’s cross-modal transferability and pinpoint critical deployment bottlenecks. Our synthesis delivers a methodological framework and reproducible implementation pathways for lightweight adaptation of general-purpose foundation models, advancing efficient, robust, and scalable model customization paradigms.
To address parameter redundancy, high inference overhead, and catastrophic forgetting in continual learning caused by stacking multiple LoRA adapters, this paper proposes Continual Low-Rank Adaptation (C-LoRA). Methodologically, C-LoRA introduces a learnable routing matrix to dynamically dispatch task-specific updates onto a shared low-rank subspace, while imposing orthogonality constraints to mitigate inter-task interference and enhance knowledge reuse. Crucially, C-LoRA unifies LoRA into a single, scalable continual adaptation framework—requiring no additional parameters to accommodate arbitrarily long task sequences. Empirically, it achieves state-of-the-art accuracy and parameter efficiency across multiple continual learning benchmarks. Theoretically, we analyze the pivotal role of the routing matrix in balancing knowledge retention and transfer, revealing its function as a structured gate for selective parameter sharing. This work establishes the first parameter-efficient, adapter-based continual learning method that maintains architectural simplicity without sacrificing expressivity or scalability.
This work addresses the parameter inefficiency of conventional LoRA-based continual learning, where each task is assigned a separate adapter despite significant low-rank redundancy across tasks. The authors propose LiteLoRA, a novel approach that leverages a plug-in gating mechanism and subspace overlap analysis to selectively reuse existing adapters or instantiate new ones in a dynamic, task-aware manner. By identifying and exploiting shared low-rank subspaces among tasks, LiteLoRA achieves competitive or superior performance compared to state-of-the-art methods on standard continual learning benchmarks, while reducing the number of active adapters by 20% to 70%, thereby substantially improving parameter efficiency.
This work investigates how low-rank adaptation (LoRA) parameters influence catastrophic forgetting during fine-tuning. We systematically merge LoRA adapter weights back into the backbone to quantitatively analyze forgetting dynamics across pretraining and downstream tasks, as well as changes in model plasticity. We identify, for the first time, a “contextual forgetting” phenomenon in Vision Transformers (ViTs)—characterized by task-dependent degradation of local features—distinct from the global forgetting observed in ResNets and unreported in prior continual learning literature. Moreover, we reveal that LoRA rank exerts a dual regulatory effect: excessively low ranks exacerbate pretrained knowledge forgetting, while excessively high ranks impair downstream adaptability. Experiments span diverse multi-task continual learning scenarios. Our findings provide theoretical foundations and principled guidelines for rank selection in efficient, sustainable visual model adaptation.
To address parameter redundancy and catastrophic forgetting across tasks in replay-free continual learning with pretrained models, this paper proposes a dual-adapter collaborative architecture. First, a task-shared adapter employs orthogonal initialization and gradient-redistributed knowledge distillation to enhance cross-task knowledge reuse. Second, a task-specific adapter leverages learnable block-wise weights for fine-grained task discrimination. Integrating LoRA with the parameter-efficient fine-tuning (PEFT) paradigm, the method avoids storing historical samples or introducing full-sized adapters per task. Evaluated on multiple benchmarks, it achieves significant improvements in classification accuracy while reducing both training and inference overhead. The approach thus balances model efficiency, scalability, and sustained generalization capability—without compromising performance or requiring additional memory for past data.
To address catastrophic forgetting in continual learning of vision-language models (VLMs), this paper proposes a dynamic-rank LoRA method that requires no reference data and modifies neither the inference architecture nor pretrained weights. The method incrementally integrates new knowledge via low-rank updates while actively enhancing—rather than degrading—pretrained semantic capabilities. Its core contributions are: (1) the first module-level importance estimation mechanism that adaptively allocates LoRA ranks to jointly optimize plasticity and stability; and (2) a reference-free, unsupervised task-boundary awareness capability. Evaluated on multiple VLM continual learning benchmarks, it achieves state-of-the-art performance: average zero-shot transfer accuracy improves by 2.3%, task-sequence accuracy decay decreases by 67%, and inference overhead remains zero.
This work addresses the lack of theoretical understanding regarding the generalization gap between Low-Rank Adaptation (LoRA) and full fine-tuning. Within a simplified linear regression framework, the study systematically compares their generalization behaviors by integrating statistical learning theory with excess risk analysis. It establishes, for the first time, a theoretical guarantee that when the discrepancy between the pre-trained model and the downstream task exhibits a low-rank structure, LoRA achieves lower excess risk than full fine-tuning in both over-parameterized and under-parameterized regimes. Empirical experiments further corroborate this finding, demonstrating a non-intuitive improvement in test accuracy under low-rank constraints and thereby revealing the intrinsic mechanism underlying LoRA’s superior generalization capability.
This study addresses the unclear mechanisms underlying catastrophic forgetting induced by LoRA in continual learning. Leveraging high-dimensional online learning theory and teacher-student solvable models, we derive asymptotically exact dynamical equations for LoRA. The analysis quantifies the trade-off between interference suppression and adaptation speed in low-rank updates, elucidating the geometric role of adapter rank and the theoretical advantages of state-dependent masking strategies. Our results demonstrate that the proposed approach significantly mitigates forgetting while preserving plasticity for new tasks. Theoretical predictions exhibit strong agreement with both numerical simulations and MNIST experiments, establishing a rigorous dynamical foundation for parameter-efficient continual learning.
This work addresses the challenge in continual learning where existing low-rank adaptation (LoRA) methods struggle to balance rapid adaptation to new tasks with the retention of previously acquired knowledge, often leading to catastrophic forgetting. To overcome this limitation, we propose ProCL, a novel framework that, for the first time, integrates the complementary learning systems theory from neuroscience into continual LoRA. ProCL introduces structured procedural memory slots and employs input-conditional attention to dynamically retrieve task-specific adapters, thereby enabling synergistic local adaptation and global knowledge accumulation. This design allows semantically similar inputs to share adapter regions, preserving future learning capacity while jointly optimizing model plasticity and stability. Extensive experiments demonstrate that ProCL significantly outperforms current continual LoRA approaches across multiple benchmarks, effectively enhancing knowledge retention and mitigating catastrophic forgetting.
This work addresses the limitation of fixed-rank constraints in parameter-efficient fine-tuning, which fail to accommodate the heterogeneous rank requirements across different layers of neural networks. The authors propose LR-LoRA, a novel approach that introduces a learnable rank mechanism within the LoRA framework, enabling differentiable and dynamic optimization of the rank for each adapter layer. This method reveals a systematic disparity in rank demands between attention and MLP layers in Transformers, thereby providing a more flexible and effective inductive bias. Experimental results demonstrate that LR-LoRA significantly outperforms existing parameter-efficient fine-tuning methods across multiple benchmarks for language understanding and commonsense reasoning, achieving state-of-the-art performance.
This work addresses the challenges of knowledge compression and inter-task interference in continual learning, which are exacerbated by energy dispersion across model parameters. To mitigate these issues, the authors propose Energy-concentrated and Ordered Low-Rank Adaptation (E²-LoRA), a novel mechanism that analyzes the low-rank structure of output feature drift. By preserving critical parameters along principal component directions to minimize reconstruction error, E²-LoRA explicitly compresses and orders accumulated knowledge into leading singular subspaces, thereby freeing model capacity for subsequent tasks. Coupled with a dynamic rank allocation strategy, the method jointly optimizes knowledge retention and model plasticity. Extensive experiments on multiple continual learning benchmarks demonstrate that E²-LoRA significantly outperforms existing approaches, achieving state-of-the-art performance.