layer-wise adapter composition

Designs and implements modular, per-layer adapter modules—including low-rank/LoRA-style adapters—and the interfaces to attach and manage them on each network layer. Builds and analyzes mechanisms for selectively activating or composing these adapters at inference or fine-tuning time without joint retraining, with controls to avoid parameter-level interference between concepts.

layer-wiseadaptercomposition

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Low-Rank Adaptation for Foundation Models: A Comprehensive Review

Dec 31, 2024
MY
Menglin Yang
🏛️ Yale University | Nanyang Technological University | The Chinese University of Hong Kong | University of Electronic Science and Technology of China | Hong Kong University | Birla Institute of Technology and Science | Logs AI

To address the high computational cost and poor generalization in efficient adaptation of foundation models, this paper presents the first systematic survey of Low-Rank Adaptation (LoRA) extensions across broad classes of foundation models—including multimodal and scientific computing models. We propose a unified taxonomy that integrates matrix low-rank decomposition, modular adapter design, gradient-constrained optimization, and cross-task transfer analysis—thereby identifying key theoretical gaps and charting a new direction toward robustness-aware modeling. Covering over 100 state-of-the-art works, we uncover common mechanisms underlying LoRA’s cross-modal transferability and pinpoint critical deployment bottlenecks. Our synthesis delivers a methodological framework and reproducible implementation pathways for lightweight adaptation of general-purpose foundation models, advancing efficient, robust, and scalable model customization paradigms.

Computational Resource ReductionEfficient Fine-tuningLarge-scale Pre-trained Models

Must-Read Papers

Most classic and influential ideas
View more

This work addresses a central challenge in parameter-efficient fine-tuning: identifying the optimal placement of adapters to achieve peak performance with minimal parameters. The authors propose PAGE, a metric based on initial gradient energy analysis, which reveals that adaptation effects are highly concentrated in the down-projection modules of shallow feed-forward networks. Leveraging this insight, they introduce DomLoRA—a method that deploys a single LoRA adapter exclusively in this dominant module. This study is the first to demonstrate the existence, architectural dependency, and task stability of such a dominant adaptation module, establishing a new paradigm for efficient fine-tuning. Experiments show that DomLoRA, using only ~0.7% of the parameters of standard LoRA, consistently outperforms it across diverse tasks—including instruction following, mathematical reasoning, code generation, and multi-turn dialogue—and further enhances the effectiveness of other LoRA variants.

adapter placementdominant adaptation modulegradient energy

Exploring Sparse Adapters for Scalable Merging of Parameter Efficient Experts

Jul 08, 2025
SY
Samin Yeasar Arnob
🏛️ McGill University | Mila | Microsoft | Université de Montréal | ServiceNow | Georgia Institute of Technology

This work addresses the challenges of multi-task adaptation and merging in parameter-efficient fine-tuning. We propose a modular architecture based on sparse adapters, integrating low-rank decomposition with structured weight sparsity to enable task-specific adapter training with minimal parameter updates—eliminating the need for downstream fine-tuning and supporting plug-and-play multi-task merging. Our key contribution is a streamlined, more efficient sparse training mechanism, both conceptually and empirically superior to LoRA and full-parameter fine-tuning. Experiments demonstrate the first successful joint merging of adapters across 20 diverse NLP tasks, yielding substantial gains in in-distribution performance. The approach consistently outperforms LoRA and full-model merging baselines in multi-task merging scenarios. However, cross-task generalization remains limited and warrants further investigation.

Comparing sparse adapters with LoRA and full fine-tuningExploring sparse adapters for scalable expert mergingInvestigating merging properties across multiple NLP tasks

LoRAverse: A Submodular Framework to Retrieve Diverse Adapters for Diffusion Models

Oct 16, 2025
MS
Mert Sonmezer
🏛️ Middle East Technical University | Virginia Tech

Facing the challenge of inefficient navigation and filtering among over 100,000 LoRA adapters hosted on large-scale platforms, this paper proposes a submodular optimization-based adapter selection framework. It formulates adapter retrieval as a combinatorial optimization problem balancing relevance and diversity. Methodologically, we design a differentiable submodular objective function that jointly incorporates low-rank decomposition features from attention layers and semantic similarity metrics, enabling end-to-end optimization. We evaluate the approach quantitatively (retrieval accuracy, diversity score) and qualitatively (generation quality, stylistic coverage) on text-to-image generation tasks. Experiments demonstrate significant improvements across multiple domains: +12.3% Recall@10 in adapter retrieval and +28.6% LPIPS diversity gain, while preserving generation fidelity. This work establishes a novel paradigm for efficient utilization of large-scale lightweight adapter repositories.

Addressing navigation challenges in massive unorganized adapter collectionsSelecting relevant diverse LoRA adapters from vast databasesSolving combinatorial optimization for optimal adapter selection

PLoP: Precise LoRA Placement for Efficient Finetuning of Large Models

Jun 25, 2025
SH
Soufiane Hayou
🏛️ Simons Institute | UC Berkeley | Flatiron Institute

In LoRA fine-tuning of large language models, adapter placement is typically determined heuristically, lacking principled theoretical guidance and requiring labor-intensive manual trial-and-error. Method: This paper proposes PLoP, a lightweight, automated LoRA placement framework grounded in gradient sensitivity analysis and module importance estimation. PLoP introduces the first theory-driven mechanism for selecting optimal adapter locations—dynamically identifying task-adaptive module types (e.g., attention or MLP sublayers) without human intervention. It incurs no additional training overhead and seamlessly integrates with standard LoRA pipelines. Contribution/Results: Evaluated on supervised fine-tuning and reinforcement learning for reasoning, PLoP consistently matches or surpasses prevalent hand-crafted strategies (e.g., full-attention or full-MLP placement), demonstrating strong effectiveness, generalizability across tasks and architectures, and plug-and-play compatibility.

Automating module selection for LoRA adaptationImproving performance over manual placement strategiesOptimizing LoRA adapter placement for efficient finetuning

This work addresses a critical limitation in existing mixture-of-LoRA models, where imbalanced routing weights often lead to the activation of only a few adapters, thereby constraining model expressiveness. To overcome this, the authors propose ReMix, a novel approach that eliminates learnable routing weights and instead introduces a non-learnable, balanced routing mechanism. By leveraging the REINFORCE leave-one-out (RLOO) gradient estimator from reinforcement learning, ReMix constructs an unbiased, non-differentiable routing policy that ensures equal contribution from all activated LoRA modules. Under the constraint of identical active parameter counts, ReMix significantly outperforms current state-of-the-art parameter-efficient fine-tuning methods, effectively breaking through the performance bottleneck inherent in conventional mixture-of-LoRA architectures.

expressive powerlow-rank adaptersMixture-of-LoRAs

Latest Papers

What's happening recently
View more

This study addresses the challenges of weight interference and retraining costs in merging multiple LoRA adapters by proposing the READ method. Built upon fixed factorization and a unidirectional coupling mechanism, READ enforces balanced canonical form constraints that enable new skills to read legacy inputs without overwriting existing outputs. This design facilitates zero-overhead skill composition during inference without requiring retraining. Experimental results demonstrate that READ surpasses baseline methods by over 20 points on benchmarks such as SuperGLUE while effectively preserving model singularity. By enabling seamless multi-skill integration within parameter-efficient fine-tuning frameworks, this work establishes a novel paradigm for adapter-based model composition.

interferenceLoRA compositionLow-rank adapters

Existing parameter-efficient fine-tuning (PEFT) methods for Mixture-of-Experts (MoE) models struggle to simultaneously leverage routing priors, enable dynamic adaptation, and facilitate cross-expert knowledge sharing, often resulting in suboptimal efficiency, high catastrophic forgetting risk, or constrained model capacity. To address these limitations, this work proposes MoE²-LoRA, the first approach that integrates the MoE paradigm into LoRA-based fine-tuning. It introduces a Routing-Conditioned Projection (RCP) module that aligns pre-trained expert specialization with task-specific adaptation and constructs a globally shared pool of LoRA experts. This design enables dynamic low-rank adaptation guided by the original router activations and promotes cross-layer knowledge sharing. Evaluated across MoE backbones of varying scales and granularities, MoE²-LoRA achieves state-of-the-art performance on downstream tasks while demonstrating superior generalization capabilities.

expert routinglow-rank adaptationMixture-of-Experts

Hot Scholars

WW

Weiping Wang

School of Information Science and Engineering, Central South University
Computer NetworkNetwork Security
YM

Yanan Ma

City University of Hong Kong
Wireless networksEdge intelligence
HR

Hossein Rahmani

Professor, Lancaster University
Computer VisionMachine LearningVideo AnalysisAction Recognition
YW

Yee Whye Teh

Professor of Statistical Machine Learning, Oxford, Research Scientist, DeepMind
Machine LearningArtificial IntelligenceStatisticsComputer Science