Score
Designs and implements modular adapter components and the adapter network architectures that host them, specifying adapter parameterizations (e.g., low-rank or weight‑decomposed LoRA) and the attachment points into a base model. Builds and evaluates composition mechanisms and control semantics — hierarchical stacking, role‑based (DORA‑RBAC) composition, and adapter network designs — that enable reusable, modular task or domain behavior without retraining the base model.
This work addresses cross-domain interference in multi-domain adapter composition for large language models by proposing the DoRA-RBAC hierarchical composition framework and systematically comparing Euclidean averaging with Riemannian geometry–based directional normalization fusion strategies. Through experiments on multiple question-answering benchmarks, the study provides the first empirical evidence that orthogonality and angular alignment of parameter updates are not reliable predictors of adapter composition performance, thereby challenging the prevailing assumption that the geometric structure of parameter space primarily governs interference. The results demonstrate that, although single-domain performance matches that of LoRA, geometry-aware fusion does not significantly outperform standard averaging in multi-domain settings, suggesting that interference likely arises from interactions within shared nonlinear representations rather than parameter-space geometry.
This work addresses a central challenge in parameter-efficient fine-tuning: identifying the optimal placement of adapters to achieve peak performance with minimal parameters. The authors propose PAGE, a metric based on initial gradient energy analysis, which reveals that adaptation effects are highly concentrated in the down-projection modules of shallow feed-forward networks. Leveraging this insight, they introduce DomLoRA—a method that deploys a single LoRA adapter exclusively in this dominant module. This study is the first to demonstrate the existence, architectural dependency, and task stability of such a dominant adaptation module, establishing a new paradigm for efficient fine-tuning. Experiments show that DomLoRA, using only ~0.7% of the parameters of standard LoRA, consistently outperforms it across diverse tasks—including instruction following, mathematical reasoning, code generation, and multi-turn dialogue—and further enhances the effectiveness of other LoRA variants.
This work addresses the challenges of multi-task adaptation and merging in parameter-efficient fine-tuning. We propose a modular architecture based on sparse adapters, integrating low-rank decomposition with structured weight sparsity to enable task-specific adapter training with minimal parameter updates—eliminating the need for downstream fine-tuning and supporting plug-and-play multi-task merging. Our key contribution is a streamlined, more efficient sparse training mechanism, both conceptually and empirically superior to LoRA and full-parameter fine-tuning. Experiments demonstrate the first successful joint merging of adapters across 20 diverse NLP tasks, yielding substantial gains in in-distribution performance. The approach consistently outperforms LoRA and full-model merging baselines in multi-task merging scenarios. However, cross-task generalization remains limited and warrants further investigation.
Large-scale pretrained models face high computational overhead and structural instability during multi-task adaptation. Method: This paper proposes a composable fine-tuning framework that integrates graph-structured task priors with modular adapters. It constructs a task-relation graph to model inter-task dependencies, leveraging this structured prior to guide low-rank adapter parameter allocation and dynamic routing. The framework incorporates plug-and-play adapter design, relation-matrix regularization, and temperature- and gating-based control mechanisms to mitigate path conflicts and redundant computation. Contributions/Results: Experiments demonstrate significant improvements in task prediction accuracy and adapter assignment precision. The method exhibits strong robustness under hyperparameter, environmental, and data perturbations, achieving both high performance and parameter efficiency. It establishes a new paradigm for multi-task adaptation—characterized by interpretability, reusability, and structural stability—without compromising scalability or practicality.
This work addresses a critical limitation in existing mixture-of-LoRA models, where imbalanced routing weights often lead to the activation of only a few adapters, thereby constraining model expressiveness. To overcome this, the authors propose ReMix, a novel approach that eliminates learnable routing weights and instead introduces a non-learnable, balanced routing mechanism. By leveraging the REINFORCE leave-one-out (RLOO) gradient estimator from reinforcement learning, ReMix constructs an unbiased, non-differentiable routing policy that ensures equal contribution from all activated LoRA modules. Under the constraint of identical active parameter counts, ReMix significantly outperforms current state-of-the-art parameter-efficient fine-tuning methods, effectively breaking through the performance bottleneck inherent in conventional mixture-of-LoRA architectures.
该研究针对MoE模型参数高效微调中适配器碎片化问题,提出ACE方法,通过整合冗余专家的适配器并执行分组计算,提高准确率和训练速度。
本文通过图论框架分析了LoRA适配器在生产环境下的共现模式,揭示了其稀疏性和重尾分布特征,并提出基于共现统计的预加载策略以优化执行延迟和资源管理。
研究通过引入Jaccard路由重叠和适配器梯度余弦相似性,诊断MoE+LoRA微调中的领域间竞争问题,并提出SpawnLoRA方法来动态增加门控子适配器以减少负面迁移。
This study addresses the challenges of weight interference and retraining costs in merging multiple LoRA adapters by proposing the READ method. Built upon fixed factorization and a unidirectional coupling mechanism, READ enforces balanced canonical form constraints that enable new skills to read legacy inputs without overwriting existing outputs. This design facilitates zero-overhead skill composition during inference without requiring retraining. Experimental results demonstrate that READ surpasses baseline methods by over 20 points on benchmarks such as SuperGLUE while effectively preserving model singularity. By enabling seamless multi-skill integration within parameter-efficient fine-tuning frameworks, this work establishes a novel paradigm for adapter-based model composition.
This work challenges the prevailing assumption that parameter-efficient fine-tuning methods such as LoRA encode only “skills” rather than memorized data, by precisely quantifying—in bits—the amount of information written into adapters while the backbone model remains frozen. Leveraging compression-based memory analysis and information-theoretic measures on the Qwen2.5 model, the study reveals that each trainable parameter stores only a few bits of information, with MLP layers exhibiting substantially higher storage efficiency than attention layers. Memory capacity is found to be governed primarily by parameter location rather than sheer parameter count. Furthermore, adapters trained via supervised learning pose notable privacy risks due to data memorization, whereas those trained through reinforcement learning retain almost no secrets from the original training data.