Score
Designs and implements low‑rank adapter modules that impose orthogonality constraints or perform orthogonal projections between update subspaces (orthogonal LoRA / orthogonal projection LoRA) to produce parameter‑efficient update matrices. Builds analyses and tooling to compute post‑training contribution coordinates and to add, combine, or delete those orthogonal projections history‑free while minimizing entanglement among updates.
This work addresses the fragmented landscape of LoRA variants, which currently lack a unified taxonomy, theoretical framework, and standardized implementation and evaluation protocols. To this end, we propose the first four-dimensional classification scheme grounded in rank structure, optimization dynamics, initialization strategies, and MoE integration, offering a cohesive theoretical perspective. We further develop LoRAFactory, a modular codebase enabling systematic experimentation across diverse tasks—including natural language generation, natural language understanding, and image classification—through large-scale empirical studies. Our findings reveal that the original LoRA, when equipped with well-tuned hyperparameters, matches or surpasses most existing variants, while exhibiting pronounced sensitivity to learning rate choices. These results underscore LoRA’s robustness and efficacy, establishing it as a standardized benchmark for parameter-efficient fine-tuning.
Existing methods struggle to effectively merge multiple task-specific LoRA adapters into a single low-rank adapter without causing capability fragmentation or violating the low-rank structure. This work proposes a novel “Compress-then-Merge” (CtM) paradigm: it first constructs a shared r-dimensional subspace from the LoRA weights and orthogonally projects each adapter onto this subspace, then performs standard merging within the resulting r×r core coordinate space, followed by truncated SVD to strictly enforce the target rank constraint. Evaluated across multiple models and tasks, CtM significantly outperforms existing single-LoRA baselines and substantially narrows the performance gap with full-parameter merging, achieving for the first time an efficient and rank-preserving LoRA fusion.
This work addresses the trade-off between performance and efficiency in parameter-efficient fine-tuning of large language models by systematically reinterpreting Low-Rank Adaptation (LoRA) through the lens of signal processing. Leveraging classical low-rank modeling and inverse problem theory, it establishes a unified framework to understand both existing and future efficient fine-tuning methods. The study proposes a three-dimensional technical framework encompassing architecture design, optimization strategies, and full-lifecycle deployment, integrating core techniques such as singular value decomposition, rank expansion, cross-layer tensorization, norm-invariant optimization, and parameterization-aware solvers. This approach provides theoretical grounding and principled design guidelines for LoRA and its variants, while extending their applicability across pre-training, post-training, and deployment stages, thereby fostering bidirectional integration between signal processing and deep learning.
To address parameter redundancy in Low-Rank Adaptation (LoRA), this paper proposes SymLoRA—a computationally efficient fine-tuning method that models adapter weights as symmetric low-rank matrices and replaces the conventional BA decomposition with a spectral decomposition Q diag(Λ) Qᵀ. Its core innovation lies in the first introduction of symmetry constraints and spectral parameterization for LoRA, coupled with an SVD-inspired initialization strategy, enabling more concise theoretical modeling and improved optimization stability in end-to-end training. Evaluated across multiple NLP benchmarks, SymLoRA reduces trainable parameters by 48%–52% compared to standard LoRA—halving parameter count—thereby significantly lowering GPU memory consumption and computational overhead. Crucially, it retains downstream task performance on par with LoRA, with no discernible accuracy degradation.
This work addresses the inefficiency in spectral utilization within the low-rank subspace of trained LoRA adapters, where many singular directions are either unhelpful or detrimental to downstream tasks. The authors propose the first training-free post-processing method for LoRA: by performing SVD on the trained LoRA weights and estimating the sensitivity of each singular component via gradients computed on a small calibration set, they reweight the singular values according to their sensitivity while preserving the original singular directions. This approach adjusts only around 1,000 scalar coefficients and yields significant performance gains across four benchmarks on Llama-3.1-8B and Qwen3-8B, achieving up to a 4.4-point improvement on CommonsenseQA and a 2.4-point gain in HumanEval pass@1.
During multi-LoRA merging, semantic vectors interfere with each other, undermining composability. A prevalent misconception equates orthogonality with semantic decoupling. Method: We propose Orthogonal Monte Carlo Dropout (OMCD), the first method to achieve *strictly orthogonal*, sparse semantic vectors—guaranteed both theoretically and at runtime—without inference overhead. OMCD integrates LoRA fine-tuning, sparsity-inducing modeling, and explicit orthogonality constraints, leveraging Monte Carlo Dropout for efficient orthogonal merging. Results: Experiments show OMCD significantly suppresses direct interference among modules. However, it reveals that orthogonality alone is insufficient for semantic composability—challenging the implicit “orthogonality implies decoupling” assumption in adapter fusion. This work provides new theoretical insights and practical guidance for designing composable adapters.
本文针对LoRA在适应大型预训练模型时的秩利用率问题,提出了一种通过联合切空间优化的方法ISO-LoRA,以提高有效秩和下游任务表现。
This study addresses the significant performance gap between Low-Rank Adaptation (LoRA) and full fine-tuning. To bridge this disparity, we propose GDLoRA, which introduces a novel orthogonal decomposition mechanism based on the reachable gradient space of LoRA. Specifically, our method extracts the orthogonal component of the gradient to directly update the base model weights. By integrating the AdamW optimizer with forward and backward signal reconstruction techniques, GDLoRA achieves effective optimization of base parameters without incurring additional memory overhead. Extensive experiments demonstrate that GDLoRA significantly outperforms standard LoRA across natural language understanding and mathematical reasoning tasks, substantially narrowing the performance gap with full fine-tuning. This work establishes a new paradigm for parameter-efficient fine-tuning.
This work addresses the lack of theoretical convergence guarantees for Low-Rank Adaptation (LoRA) in stochastic optimization, particularly regarding the query complexity required to reach an ε-stationary point. To bridge this gap, the authors introduce LoRA-STORM, the first variance-reduced algorithm tailored for LoRA, incorporating momentum and leveraging a mean-squared smoothness assumption to establish tighter convergence bounds. Under deterministic settings, LoRA-GD achieves a gradient complexity of O(ε⁻⁴). In stochastic settings, the proposed methods LoRA-NSGDM and LoRA-STORM attain stochastic oracle complexities of O(ε⁻⁸) and O(ε⁻⁶), respectively, substantially improving upon existing results in the literature.
This study addresses the issue in LoRA fine-tuning where a few singular directions dominate updates, causing adaptation imbalance and degrading general model capabilities. To mitigate this, it proposes LoRA-Norm, a post-training normalization method that integrates spectral rebalancing with nuclear norm recovery. This approach reveals that balancing adapter gain and equalizing responses constitute independent optimization objectives, enabling gain rebalancing while preserving effective learning directions. Notably, the method requires no additional data or training overhead, achieving optimization at zero inference cost. Extensive evaluations across diverse backbone architectures and tasks demonstrate that LoRA-Norm significantly improves both average specialization and capability retention, outperforming existing spectral pruning and gradient editing techniques.
This study addresses the prevalent misconception that the rank parameter in LoRA functions as a capacity control mechanism, clarifying its true role under norm constraints. By integrating spectral analysis, distribution distance metrics, and Rademacher complexity theory, this work demonstrates that under a hard norm budget, the rank upper bound becomes inactive within the nuclear norm ball. Consequently, it redefines rank as a controlling factor for the reachable update set and spectral alignment cost. The primary contributions include establishing a rank-independent generalization complexity upper bound and deriving tight bounds on the minimum rank required for source-target task alignment. These findings provide a rigorous theoretical foundation for LoRA hyperparameter selection.