Score
Designs, implements, and evaluates parameter-efficient low-rank adapter modules and their training procedures that decompose neural network weight updates into low-rank factors for fine-tuning; this includes deterministic LoRA-style adapters, sparse and variational Bayesian formulations, and decomposed expert/mixture variants. Builds methods for adapter-based incremental and few-shot tuning and analyzes rank selection, regularization, uncertainty quantification, interpretability of adapter directions, and impacts on downstream accuracy and reasoning.
This work addresses the fragmented landscape of LoRA variants, which currently lack a unified taxonomy, theoretical framework, and standardized implementation and evaluation protocols. To this end, we propose the first four-dimensional classification scheme grounded in rank structure, optimization dynamics, initialization strategies, and MoE integration, offering a cohesive theoretical perspective. We further develop LoRAFactory, a modular codebase enabling systematic experimentation across diverse tasks—including natural language generation, natural language understanding, and image classification—through large-scale empirical studies. Our findings reveal that the original LoRA, when equipped with well-tuned hyperparameters, matches or surpasses most existing variants, while exhibiting pronounced sensitivity to learning rate choices. These results underscore LoRA’s robustness and efficacy, establishing it as a standardized benchmark for parameter-efficient fine-tuning.
To address the high computational cost and poor generalization in efficient adaptation of foundation models, this paper presents the first systematic survey of Low-Rank Adaptation (LoRA) extensions across broad classes of foundation models—including multimodal and scientific computing models. We propose a unified taxonomy that integrates matrix low-rank decomposition, modular adapter design, gradient-constrained optimization, and cross-task transfer analysis—thereby identifying key theoretical gaps and charting a new direction toward robustness-aware modeling. Covering over 100 state-of-the-art works, we uncover common mechanisms underlying LoRA’s cross-modal transferability and pinpoint critical deployment bottlenecks. Our synthesis delivers a methodological framework and reproducible implementation pathways for lightweight adaptation of general-purpose foundation models, advancing efficient, robust, and scalable model customization paradigms.
To address parameter redundancy in Low-Rank Adaptation (LoRA), this paper proposes SymLoRA—a computationally efficient fine-tuning method that models adapter weights as symmetric low-rank matrices and replaces the conventional BA decomposition with a spectral decomposition Q diag(Λ) Qᵀ. Its core innovation lies in the first introduction of symmetry constraints and spectral parameterization for LoRA, coupled with an SVD-inspired initialization strategy, enabling more concise theoretical modeling and improved optimization stability in end-to-end training. Evaluated across multiple NLP benchmarks, SymLoRA reduces trainable parameters by 48%–52% compared to standard LoRA—halving parameter count—thereby significantly lowering GPU memory consumption and computational overhead. Crucially, it retains downstream task performance on par with LoRA, with no discernible accuracy degradation.
This work addresses the trade-off between performance and efficiency in parameter-efficient fine-tuning of large language models by systematically reinterpreting Low-Rank Adaptation (LoRA) through the lens of signal processing. Leveraging classical low-rank modeling and inverse problem theory, it establishes a unified framework to understand both existing and future efficient fine-tuning methods. The study proposes a three-dimensional technical framework encompassing architecture design, optimization strategies, and full-lifecycle deployment, integrating core techniques such as singular value decomposition, rank expansion, cross-layer tensorization, norm-invariant optimization, and parameterization-aware solvers. This approach provides theoretical grounding and principled design guidelines for LoRA and its variants, while extending their applicability across pre-training, post-training, and deployment stages, thereby fostering bidirectional integration between signal processing and deep learning.
This work addresses the challenges of multi-task adaptation and merging in parameter-efficient fine-tuning. We propose a modular architecture based on sparse adapters, integrating low-rank decomposition with structured weight sparsity to enable task-specific adapter training with minimal parameter updates—eliminating the need for downstream fine-tuning and supporting plug-and-play multi-task merging. Our key contribution is a streamlined, more efficient sparse training mechanism, both conceptually and empirically superior to LoRA and full-parameter fine-tuning. Experiments demonstrate the first successful joint merging of adapters across 20 diverse NLP tasks, yielding substantial gains in in-distribution performance. The approach consistently outperforms LoRA and full-model merging baselines in multi-task merging scenarios. However, cross-task generalization remains limited and warrants further investigation.
This work addresses the limitation of fixed-rank constraints in parameter-efficient fine-tuning, which fail to accommodate the heterogeneous rank requirements across different layers of neural networks. The authors propose LR-LoRA, a novel approach that introduces a learnable rank mechanism within the LoRA framework, enabling differentiable and dynamic optimization of the rank for each adapter layer. This method reveals a systematic disparity in rank demands between attention and MLP layers in Transformers, thereby providing a more flexible and effective inductive bias. Experimental results demonstrate that LR-LoRA significantly outperforms existing parameter-efficient fine-tuning methods across multiple benchmarks for language understanding and commonsense reasoning, achieving state-of-the-art performance.
Standard LoRA lacks the ability to quantify uncertainty, often leading to overconfident and poorly calibrated models. While existing Bayesian LoRA variants introduce uncertainty modeling, they suffer from parameter inflation that causes training instability. This work proposes a novel parameter-efficient Bayesian fine-tuning framework that models weight uncertainty within an extremely low-dimensional projected subspace. It reveals for the first time that the weight covariance of large models exhibits a low-rank structure, enabling stable and efficient Bayesian optimization. By leveraging this insight, the method significantly improves model calibration and generalization with negligible parameter overhead, demonstrating the effectiveness and scalability of uncertainty modeling in low-dimensional subspaces.
This work addresses the parameter inefficiency of conventional LoRA-based continual learning, where each task is assigned a separate adapter despite significant low-rank redundancy across tasks. The authors propose LiteLoRA, a novel approach that leverages a plug-in gating mechanism and subspace overlap analysis to selectively reuse existing adapters or instantiate new ones in a dynamic, task-aware manner. By identifying and exploiting shared low-rank subspaces among tasks, LiteLoRA achieves competitive or superior performance compared to state-of-the-art methods on standard continual learning benchmarks, while reducing the number of active adapters by 20% to 70%, thereby substantially improving parameter efficiency.
该研究针对MoE模型参数高效微调中适配器碎片化问题,提出ACE方法,通过整合冗余专家的适配器并执行分组计算,提高准确率和训练速度。
This study addresses the long-standing absence of theoretical guidance for selecting adapter ranks in LoRA fine-tuning, which complicates the trade-off between optimization conditioning and computational efficiency. Drawing on non-convex low-rank matrix sensing theory, this work proposes a data-dependent LoRA-RIP metric to characterize the optimization geometry of cross-entropy objectives. It further demonstrates that sufficient overparameterization eliminates spurious local minima, surpassing the classical 1/3 theoretical guarantee. Crucially, the analysis reveals that rank and data quality constitute coupled resources requiring joint consideration. Experiments across language and vision tasks validate these theoretical predictions, indicating that the proposed framework substantially enhances both the efficiency and reliability of parameter-efficient fine-tuning.
This study addresses the significant performance gap between Low-Rank Adaptation (LoRA) and full fine-tuning. To bridge this disparity, we propose GDLoRA, which introduces a novel orthogonal decomposition mechanism based on the reachable gradient space of LoRA. Specifically, our method extracts the orthogonal component of the gradient to directly update the base model weights. By integrating the AdamW optimizer with forward and backward signal reconstruction techniques, GDLoRA achieves effective optimization of base parameters without incurring additional memory overhead. Extensive experiments demonstrate that GDLoRA significantly outperforms standard LoRA across natural language understanding and mathematical reasoning tasks, substantially narrowing the performance gap with full fine-tuning. This work establishes a new paradigm for parameter-efficient fine-tuning.
This work investigates whether low-rank adaptation (LoRA) can maintain strong generalization under structural constraints and proposes a more efficient, sparsity-aware fine-tuning approach. To this end, we introduce Cheap LoRA (cLA), along with its stochastic and cyclic chain variants, which enhance LoRA with sparsity. Theoretically, we establish the first information-theoretic generalization error bound for sparse LoRA and formalize cLA as a structured instance of asymmetric LoRA. Methodologically, our approach integrates sparse low-rank decomposition, randomly fixed-factor training, and cyclic parameterization. Extensive experiments across 10 pretrained models and 14 datasets demonstrate that cLA matches the performance of standard LoRA at equivalent parameter counts while reducing training time by up to 10% and peak GPU memory consumption by as much as 15%.