Score
Design, build, or analyze low‑rank adaptation (LoRA) modules whose parameters are conditioned and applied per image or multimodal context to adapt large visual or vision‑language models; this includes methods for per‑image modulation, multi‑sparsity and progressive/step‑wise adapter scheduling, cross‑contextual or dual‑stream fusion of adapters, and guidance‑driven adjustments (e.g., reflection severity) to suppress artifacts while preserving fine details.
This work addresses the fragmented landscape of LoRA variants, which currently lack a unified taxonomy, theoretical framework, and standardized implementation and evaluation protocols. To this end, we propose the first four-dimensional classification scheme grounded in rank structure, optimization dynamics, initialization strategies, and MoE integration, offering a cohesive theoretical perspective. We further develop LoRAFactory, a modular codebase enabling systematic experimentation across diverse tasks—including natural language generation, natural language understanding, and image classification—through large-scale empirical studies. Our findings reveal that the original LoRA, when equipped with well-tuned hyperparameters, matches or surpasses most existing variants, while exhibiting pronounced sensitivity to learning rate choices. These results underscore LoRA’s robustness and efficacy, establishing it as a standardized benchmark for parameter-efficient fine-tuning.
To address the high computational cost and poor generalization in efficient adaptation of foundation models, this paper presents the first systematic survey of Low-Rank Adaptation (LoRA) extensions across broad classes of foundation models—including multimodal and scientific computing models. We propose a unified taxonomy that integrates matrix low-rank decomposition, modular adapter design, gradient-constrained optimization, and cross-task transfer analysis—thereby identifying key theoretical gaps and charting a new direction toward robustness-aware modeling. Covering over 100 state-of-the-art works, we uncover common mechanisms underlying LoRA’s cross-modal transferability and pinpoint critical deployment bottlenecks. Our synthesis delivers a methodological framework and reproducible implementation pathways for lightweight adaptation of general-purpose foundation models, advancing efficient, robust, and scalable model customization paradigms.
Existing LoRA adapters incur substantial computational overhead and high latency during inference serving, hindering their practical deployment in vision tasks. This paper introduces VaLoRA, an end-to-end system for efficient inference optimization of large vision-language models (LVLMs). Our approach addresses three core challenges: (1) accuracy-aware automatic generation of LoRA adapters; (2) adaptive block-wise batching supporting heterogeneous adapters; and (3) a flexible adapter orchestration mechanism integrating request-level features and domain knowledge. VaLoRA holistically integrates LoRA fine-tuning, dynamic batching, and request–adapter co-scheduling. Evaluated across five representative vision tasks and three mainstream LVLMs, VaLoRA achieves average accuracy gains of 24–62% and end-to-end latency reductions of 20–89% over baseline methods.
To address the limited generalization of Low-Rank Adaptation (LoRA) in multimodal tasks, this paper proposes MoSLoRA—a novel subspace-aware LoRA variant. MoSLoRA reformulates LoRA by decomposing the weight matrix into two orthogonal subspaces and introducing a learnable linear mixer for dynamic fusion. Crucially, it jointly optimizes both subspace representations and mixing coefficients without increasing inference overhead and remains fully modality-agnostic. Extensive experiments across three diverse multimodal tasks—commonsense reasoning, vision-instruction fine-tuning, and topic-driven text-to-image generation—demonstrate consistent and significant improvements over standard LoRA, with substantial average gains. These results validate MoSLoRA’s effectiveness and cross-modal robustness. The core contribution lies in the synergistic design of subspace decoupling and learnable mixing, establishing a new paradigm for parameter-efficient fine-tuning.
In LoRA fine-tuning of large language models, adapter placement is typically determined heuristically, lacking principled theoretical guidance and requiring labor-intensive manual trial-and-error. Method: This paper proposes PLoP, a lightweight, automated LoRA placement framework grounded in gradient sensitivity analysis and module importance estimation. PLoP introduces the first theory-driven mechanism for selecting optimal adapter locations—dynamically identifying task-adaptive module types (e.g., attention or MLP sublayers) without human intervention. It incurs no additional training overhead and seamlessly integrates with standard LoRA pipelines. Contribution/Results: Evaluated on supervised fine-tuning and reinforcement learning for reasoning, PLoP consistently matches or surpasses prevalent hand-crafted strategies (e.g., full-attention or full-MLP placement), demonstrating strong effectiveness, generalizability across tasks and architectures, and plug-and-play compatibility.
While LoRA is parameter-efficient, flat solutions in its low-rank optimization subspace may still correspond to sharp directions in the full-parameter space, harming generalization. Method: This paper introduces Bayesian-LoRA—the first approach to explicitly incorporate loss surface flatness constraints into the LoRA objective. It designs a lightweight stochastic weight perturbation scheme grounded in Bayesian expected loss, avoiding high-overhead second-order methods (e.g., SAM) and eliminating the need for extra backpropagation or Hessian computation. Perturbation and low-rank decomposition are jointly optimized. Results: Bayesian-LoRA significantly improves generalization across diverse NLP and image classification tasks and architectures, with training cost comparable to standard LoRA. Its core contribution is establishing the first theoretical link between LoRA optimization and full-parameter-space flatness, yielding the first gradient-free, perturbation-based PEFT method that simultaneously achieves efficiency and strong generalization.
To address harmful redundant parameters, catastrophic forgetting of general knowledge, and degraded downstream performance in vision-instruction fine-tuning of multimodal large language models (MLLMs) using LoRA, this paper proposes a synergistic framework combining sparse parameter updates with conflict-mitigating regularization. We introduce, for the first time within the LoRA paradigm, a theoretically grounded structured sparsity mechanism for parameter updates, alongside a knowledge-conflict-aware regularizer that explicitly suppresses interference between general and task-specific knowledge at the update-trajectory level. Our method improves both general capabilities (MMMU ↑) and downstream performance (OCRBench ↑), while adding ≤5% trainable parameters—outperforming standard LoRA and other adaptation methods. It effectively mitigates catastrophic forgetting and achieves balanced optimization of generality and specialization.
Standard LoRA faces limitations in parameter-efficient fine-tuning due to the difficulty of predefining the optimal rank, sensitivity to hyperparameters, and complexity in heterogeneous deployment. This work proposes first training LoRA modules at a high rank and then applying post-training compression or dynamic rank annealing during training via weight update matrix reconstruction combined with randomized singular value decomposition (RSVD), thereby overcoming the expressivity bottleneck of direct low-rank training. The approach significantly outperforms standard LoRA trained directly at the same target rank across 13 text and 10 vision-language tasks, with particularly pronounced gains at extremely low target ranks, achieving a superior performance–parameter trade-off.
This work addresses the trade-off between performance and efficiency in parameter-efficient fine-tuning of large language models by systematically reinterpreting Low-Rank Adaptation (LoRA) through the lens of signal processing. Leveraging classical low-rank modeling and inverse problem theory, it establishes a unified framework to understand both existing and future efficient fine-tuning methods. The study proposes a three-dimensional technical framework encompassing architecture design, optimization strategies, and full-lifecycle deployment, integrating core techniques such as singular value decomposition, rank expansion, cross-layer tensorization, norm-invariant optimization, and parameterization-aware solvers. This approach provides theoretical grounding and principled design guidelines for LoRA and its variants, while extending their applicability across pre-training, post-training, and deployment stages, thereby fostering bidirectional integration between signal processing and deep learning.
This work investigates whether low-rank adaptation (LoRA) can maintain strong generalization under structural constraints and proposes a more efficient, sparsity-aware fine-tuning approach. To this end, we introduce Cheap LoRA (cLA), along with its stochastic and cyclic chain variants, which enhance LoRA with sparsity. Theoretically, we establish the first information-theoretic generalization error bound for sparse LoRA and formalize cLA as a structured instance of asymmetric LoRA. Methodologically, our approach integrates sparse low-rank decomposition, randomly fixed-factor training, and cyclic parameterization. Extensive experiments across 10 pretrained models and 14 datasets demonstrate that cLA matches the performance of standard LoRA at equivalent parameter counts while reducing training time by up to 10% and peak GPU memory consumption by as much as 15%.
This work addresses a central challenge in parameter-efficient fine-tuning: identifying the optimal placement of adapters to achieve peak performance with minimal parameters. The authors propose PAGE, a metric based on initial gradient energy analysis, which reveals that adaptation effects are highly concentrated in the down-projection modules of shallow feed-forward networks. Leveraging this insight, they introduce DomLoRA—a method that deploys a single LoRA adapter exclusively in this dominant module. This study is the first to demonstrate the existence, architectural dependency, and task stability of such a dominant adaptation module, establishing a new paradigm for efficient fine-tuning. Experiments show that DomLoRA, using only ~0.7% of the parameters of standard LoRA, consistently outperforms it across diverse tasks—including instruction following, mathematical reasoning, code generation, and multi-turn dialogue—and further enhances the effectiveness of other LoRA variants.
This work addresses the limitations of LoRA fine-tuning in personalized image generation, where fixed-rank adaptation struggles to balance performance and memory efficiency and fails to accommodate varying subject complexities. The authors propose LoRA², the first method to incorporate a variable-rank mechanism into LoRA. Leveraging a variational-inspired importance-ranking strategy, LoRA² dynamically assigns adaptive ranks to different layers during fine-tuning, allowing each layer’s rank to evolve freely according to its contribution. This approach overcomes the rigidity of conventional fixed-rank designs. Experimental results across 29 subjects demonstrate that LoRA² outperforms high-rank LoRA, achieving a superior trade-off among DINO, CLIP-I, and CLIP-T metrics while significantly reducing both memory consumption and average rank size.