adapter composition

Designs and implements modular adapter components and the adapter network architectures that host them, specifying adapter parameterizations (e.g., low-rank or weight‑decomposed LoRA) and the attachment points into a base model. Builds and evaluates composition mechanisms and control semantics — hierarchical stacking, role‑based (DORA‑RBAC) composition, and adapter network designs — that enable reusable, modular task or domain behavior without retraining the base model.

adaptercomposition

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$199K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses cross-domain interference in multi-domain adapter composition for large language models by proposing the DoRA-RBAC hierarchical composition framework and systematically comparing Euclidean averaging with Riemannian geometry–based directional normalization fusion strategies. Through experiments on multiple question-answering benchmarks, the study provides the first empirical evidence that orthogonality and angular alignment of parameter updates are not reliable predictors of adapter composition performance, thereby challenging the prevailing assumption that the geometric structure of parameter space primarily governs interference. The results demonstrate that, although single-domain performance matches that of LoRA, geometry-aware fusion does not significantly outperform standard averaging in multi-domain settings, suggesting that interference likely arises from interactions within shared nonlinear representations rather than parameter-space geometry.

adapter interferencelanguage modelsmodular access control

This work addresses a central challenge in parameter-efficient fine-tuning: identifying the optimal placement of adapters to achieve peak performance with minimal parameters. The authors propose PAGE, a metric based on initial gradient energy analysis, which reveals that adaptation effects are highly concentrated in the down-projection modules of shallow feed-forward networks. Leveraging this insight, they introduce DomLoRA—a method that deploys a single LoRA adapter exclusively in this dominant module. This study is the first to demonstrate the existence, architectural dependency, and task stability of such a dominant adaptation module, establishing a new paradigm for efficient fine-tuning. Experiments show that DomLoRA, using only ~0.7% of the parameters of standard LoRA, consistently outperforms it across diverse tasks—including instruction following, mathematical reasoning, code generation, and multi-turn dialogue—and further enhances the effectiveness of other LoRA variants.

adapter placementdominant adaptation modulegradient energy

Exploring Sparse Adapters for Scalable Merging of Parameter Efficient Experts

Jul 08, 2025
SY
Samin Yeasar Arnob
🏛️ McGill University | Mila | Microsoft | Université de Montréal | ServiceNow | Georgia Institute of Technology

This work addresses the challenges of multi-task adaptation and merging in parameter-efficient fine-tuning. We propose a modular architecture based on sparse adapters, integrating low-rank decomposition with structured weight sparsity to enable task-specific adapter training with minimal parameter updates—eliminating the need for downstream fine-tuning and supporting plug-and-play multi-task merging. Our key contribution is a streamlined, more efficient sparse training mechanism, both conceptually and empirically superior to LoRA and full-parameter fine-tuning. Experiments demonstrate the first successful joint merging of adapters across 20 diverse NLP tasks, yielding substantial gains in in-distribution performance. The approach consistently outperforms LoRA and full-model merging baselines in multi-task merging scenarios. However, cross-task generalization remains limited and warrants further investigation.

Comparing sparse adapters with LoRA and full fine-tuningExploring sparse adapters for scalable expert mergingInvestigating merging properties across multiple NLP tasks

Structural Priors and Modular Adapters in the Composable Fine-Tuning Algorithm of Large-Scale Models

Nov 06, 2025
YW
Yuxiao Wang
🏛️ University of Pennsylvania | Washington University in St. Louis | Stevens Institute of Technology | University of Southern California

Large-scale pretrained models face high computational overhead and structural instability during multi-task adaptation. Method: This paper proposes a composable fine-tuning framework that integrates graph-structured task priors with modular adapters. It constructs a task-relation graph to model inter-task dependencies, leveraging this structured prior to guide low-rank adapter parameter allocation and dynamic routing. The framework incorporates plug-and-play adapter design, relation-matrix regularization, and temperature- and gating-based control mechanisms to mitigate path conflicts and redundant computation. Contributions/Results: Experiments demonstrate significant improvements in task prediction accuracy and adapter assignment precision. The method exhibits strong robustness under hyperparameter, environmental, and data perturbations, achieving both high performance and parameter efficiency. It establishes a new paradigm for multi-task adaptation—characterized by interpretability, reusability, and structural stability—without compromising scalability or practicality.

Addressing structural instability through graph-based dependency modelingImproving parameter efficiency and training stability via modular adaptersReducing computational costs in multi-task adaptation of large-scale models

This work addresses a critical limitation in existing mixture-of-LoRA models, where imbalanced routing weights often lead to the activation of only a few adapters, thereby constraining model expressiveness. To overcome this, the authors propose ReMix, a novel approach that eliminates learnable routing weights and instead introduces a non-learnable, balanced routing mechanism. By leveraging the REINFORCE leave-one-out (RLOO) gradient estimator from reinforcement learning, ReMix constructs an unbiased, non-differentiable routing policy that ensures equal contribution from all activated LoRA modules. Under the constraint of identical active parameter counts, ReMix significantly outperforms current state-of-the-art parameter-efficient fine-tuning methods, effectively breaking through the performance bottleneck inherent in conventional mixture-of-LoRA architectures.

expressive powerlow-rank adaptersMixture-of-LoRAs

Latest Papers

What's happening recently
View more

This study addresses the challenges of weight interference and retraining costs in merging multiple LoRA adapters by proposing the READ method. Built upon fixed factorization and a unidirectional coupling mechanism, READ enforces balanced canonical form constraints that enable new skills to read legacy inputs without overwriting existing outputs. This design facilitates zero-overhead skill composition during inference without requiring retraining. Experimental results demonstrate that READ surpasses baseline methods by over 20 points on benchmarks such as SuperGLUE while effectively preserving model singularity. By enabling seamless multi-skill integration within parameter-efficient fine-tuning frameworks, this work establishes a novel paradigm for adapter-based model composition.

interferenceLoRA compositionLow-rank adapters

This work challenges the prevailing assumption that parameter-efficient fine-tuning methods such as LoRA encode only “skills” rather than memorized data, by precisely quantifying—in bits—the amount of information written into adapters while the backbone model remains frozen. Leveraging compression-based memory analysis and information-theoretic measures on the Qwen2.5 model, the study reveals that each trainable parameter stores only a few bits of information, with MLP layers exhibiting substantially higher storage efficiency than attention layers. Memory capacity is found to be governed primarily by parameter location rather than sheer parameter count. Furthermore, adapters trained via supervised learning pose notable privacy risks due to data memorization, whereas those trained through reinforcement learning retain almost no secrets from the original training data.

frozen-base modelsinformation capacityLoRA adapter

Hot Scholars

AP

Ali Payani

Cisco Systems, Georgia Tech
Natural Language ProcessingLogic and ReasoningData Efficient AIFederated Learning
TW

Tianxin Wei

University of Illinois Urbana Champaign
Trustworthy Machine LearningLLMInformation Retrieval
BD

Bo Du

Department of Management, Griffith Business School
Sustainable TransportTravel BehaviourUrban Data AnalyticsLogistics and Supply Chain
SS

Shiguang Shan

Professor of Institute of Computing Technology, Chinese Academy of Sciences
Computer VisionPattern RecognitionMachine LearningFace Recognition
MD

Mengnan Du

Assistant Professor, New Jersey Institute of Technology
ExplainabilityNatural Language ProcessingTrustworthy AI