Score
Designs and implements modular, per-layer adapter modules—including low-rank/LoRA-style adapters—and the interfaces to attach and manage them on each network layer. Builds and analyzes mechanisms for selectively activating or composing these adapters at inference or fine-tuning time without joint retraining, with controls to avoid parameter-level interference between concepts.
This work addresses the fragmented landscape of LoRA variants, which currently lack a unified taxonomy, theoretical framework, and standardized implementation and evaluation protocols. To this end, we propose the first four-dimensional classification scheme grounded in rank structure, optimization dynamics, initialization strategies, and MoE integration, offering a cohesive theoretical perspective. We further develop LoRAFactory, a modular codebase enabling systematic experimentation across diverse tasks—including natural language generation, natural language understanding, and image classification—through large-scale empirical studies. Our findings reveal that the original LoRA, when equipped with well-tuned hyperparameters, matches or surpasses most existing variants, while exhibiting pronounced sensitivity to learning rate choices. These results underscore LoRA’s robustness and efficacy, establishing it as a standardized benchmark for parameter-efficient fine-tuning.
To address the high computational cost and poor generalization in efficient adaptation of foundation models, this paper presents the first systematic survey of Low-Rank Adaptation (LoRA) extensions across broad classes of foundation models—including multimodal and scientific computing models. We propose a unified taxonomy that integrates matrix low-rank decomposition, modular adapter design, gradient-constrained optimization, and cross-task transfer analysis—thereby identifying key theoretical gaps and charting a new direction toward robustness-aware modeling. Covering over 100 state-of-the-art works, we uncover common mechanisms underlying LoRA’s cross-modal transferability and pinpoint critical deployment bottlenecks. Our synthesis delivers a methodological framework and reproducible implementation pathways for lightweight adaptation of general-purpose foundation models, advancing efficient, robust, and scalable model customization paradigms.
This work addresses a central challenge in parameter-efficient fine-tuning: identifying the optimal placement of adapters to achieve peak performance with minimal parameters. The authors propose PAGE, a metric based on initial gradient energy analysis, which reveals that adaptation effects are highly concentrated in the down-projection modules of shallow feed-forward networks. Leveraging this insight, they introduce DomLoRA—a method that deploys a single LoRA adapter exclusively in this dominant module. This study is the first to demonstrate the existence, architectural dependency, and task stability of such a dominant adaptation module, establishing a new paradigm for efficient fine-tuning. Experiments show that DomLoRA, using only ~0.7% of the parameters of standard LoRA, consistently outperforms it across diverse tasks—including instruction following, mathematical reasoning, code generation, and multi-turn dialogue—and further enhances the effectiveness of other LoRA variants.
This work addresses the challenges of multi-task adaptation and merging in parameter-efficient fine-tuning. We propose a modular architecture based on sparse adapters, integrating low-rank decomposition with structured weight sparsity to enable task-specific adapter training with minimal parameter updates—eliminating the need for downstream fine-tuning and supporting plug-and-play multi-task merging. Our key contribution is a streamlined, more efficient sparse training mechanism, both conceptually and empirically superior to LoRA and full-parameter fine-tuning. Experiments demonstrate the first successful joint merging of adapters across 20 diverse NLP tasks, yielding substantial gains in in-distribution performance. The approach consistently outperforms LoRA and full-model merging baselines in multi-task merging scenarios. However, cross-task generalization remains limited and warrants further investigation.
Facing the challenge of inefficient navigation and filtering among over 100,000 LoRA adapters hosted on large-scale platforms, this paper proposes a submodular optimization-based adapter selection framework. It formulates adapter retrieval as a combinatorial optimization problem balancing relevance and diversity. Methodologically, we design a differentiable submodular objective function that jointly incorporates low-rank decomposition features from attention layers and semantic similarity metrics, enabling end-to-end optimization. We evaluate the approach quantitatively (retrieval accuracy, diversity score) and qualitatively (generation quality, stylistic coverage) on text-to-image generation tasks. Experiments demonstrate significant improvements across multiple domains: +12.3% Recall@10 in adapter retrieval and +28.6% LPIPS diversity gain, while preserving generation fidelity. This work establishes a novel paradigm for efficient utilization of large-scale lightweight adapter repositories.
In LoRA fine-tuning of large language models, adapter placement is typically determined heuristically, lacking principled theoretical guidance and requiring labor-intensive manual trial-and-error. Method: This paper proposes PLoP, a lightweight, automated LoRA placement framework grounded in gradient sensitivity analysis and module importance estimation. PLoP introduces the first theory-driven mechanism for selecting optimal adapter locations—dynamically identifying task-adaptive module types (e.g., attention or MLP sublayers) without human intervention. It incurs no additional training overhead and seamlessly integrates with standard LoRA pipelines. Contribution/Results: Evaluated on supervised fine-tuning and reinforcement learning for reasoning, PLoP consistently matches or surpasses prevalent hand-crafted strategies (e.g., full-attention or full-MLP placement), demonstrating strong effectiveness, generalizability across tasks and architectures, and plug-and-play compatibility.
This work addresses a critical limitation in existing mixture-of-LoRA models, where imbalanced routing weights often lead to the activation of only a few adapters, thereby constraining model expressiveness. To overcome this, the authors propose ReMix, a novel approach that eliminates learnable routing weights and instead introduces a non-learnable, balanced routing mechanism. By leveraging the REINFORCE leave-one-out (RLOO) gradient estimator from reinforcement learning, ReMix constructs an unbiased, non-differentiable routing policy that ensures equal contribution from all activated LoRA modules. Under the constraint of identical active parameter counts, ReMix significantly outperforms current state-of-the-art parameter-efficient fine-tuning methods, effectively breaking through the performance bottleneck inherent in conventional mixture-of-LoRA architectures.
该研究针对MoE模型参数高效微调中适配器碎片化问题,提出ACE方法,通过整合冗余专家的适配器并执行分组计算,提高准确率和训练速度。
研究通过引入Jaccard路由重叠和适配器梯度余弦相似性,诊断MoE+LoRA微调中的领域间竞争问题,并提出SpawnLoRA方法来动态增加门控子适配器以减少负面迁移。
This study addresses the challenges of weight interference and retraining costs in merging multiple LoRA adapters by proposing the READ method. Built upon fixed factorization and a unidirectional coupling mechanism, READ enforces balanced canonical form constraints that enable new skills to read legacy inputs without overwriting existing outputs. This design facilitates zero-overhead skill composition during inference without requiring retraining. Experimental results demonstrate that READ surpasses baseline methods by over 20 points on benchmarks such as SuperGLUE while effectively preserving model singularity. By enabling seamless multi-skill integration within parameter-efficient fine-tuning frameworks, this work establishes a novel paradigm for adapter-based model composition.
Existing parameter-efficient fine-tuning (PEFT) methods for Mixture-of-Experts (MoE) models struggle to simultaneously leverage routing priors, enable dynamic adaptation, and facilitate cross-expert knowledge sharing, often resulting in suboptimal efficiency, high catastrophic forgetting risk, or constrained model capacity. To address these limitations, this work proposes MoE²-LoRA, the first approach that integrates the MoE paradigm into LoRA-based fine-tuning. It introduces a Routing-Conditioned Projection (RCP) module that aligns pre-trained expert specialization with task-specific adaptation and constructs a globally shared pool of LoRA experts. This design enables dynamic low-rank adaptation guided by the original router activations and promotes cross-layer knowledge sharing. Evaluated across MoE backbones of varying scales and granularities, MoE²-LoRA achieves state-of-the-art performance on downstream tasks while demonstrating superior generalization capabilities.
本文通过图论框架分析了LoRA适配器在生产环境下的共现模式,揭示了其稀疏性和重尾分布特征,并提出基于共现统计的预加载策略以优化执行延迟和资源管理。