Score
Designs and implements lightweight adapter modules added as auxiliary side stems or input paths to a frozen pretrained backbone, along with the tuning procedures that update only those adapter parameters (often low-rank) to adapt the model to a new task. This includes building the side-stem architecture, attention or gating mechanisms for combining stem and backbone activations, and measuring parameter-efficiency versus task performance.
Large-scale pretrained models face high computational overhead and structural instability during multi-task adaptation. Method: This paper proposes a composable fine-tuning framework that integrates graph-structured task priors with modular adapters. It constructs a task-relation graph to model inter-task dependencies, leveraging this structured prior to guide low-rank adapter parameter allocation and dynamic routing. The framework incorporates plug-and-play adapter design, relation-matrix regularization, and temperature- and gating-based control mechanisms to mitigate path conflicts and redundant computation. Contributions/Results: Experiments demonstrate significant improvements in task prediction accuracy and adapter assignment precision. The method exhibits strong robustness under hyperparameter, environmental, and data perturbations, achieving both high performance and parameter efficiency. It establishes a new paradigm for multi-task adaptation—characterized by interpretability, reusability, and structural stability—without compromising scalability or practicality.
Existing parameter-efficient fine-tuning (PEFT) methods—such as LoRA—rely on heuristic adapter architectures, suffering from poor generalization and limited transferability across models. Method: We propose the first learnable rectangular adapter search framework grounded in Monarch matrices—the first application of Monarch structure to PEFT—supported by theoretical analysis demonstrating superior expressivity over LoRA. Our approach employs differentiable neural architecture search to automatically discover optimal lightweight adapter topologies, eliminating manual specification of rank or module shape, and integrates low-parameter adapter design with efficient fine-tuning strategies. Contribution/Results: On multi-task and multi-model benchmarks, our method significantly outperforms state-of-the-art PEFT approaches, achieving comparable or superior performance using only 5% of LoRA’s parameters. It delivers both strong cross-task/model generalization and exceptional parameter efficiency.
Existing adapter-based fine-tuning methods, though parameter-efficient, suffer from substantial memory and computational overhead, prolonged training time, and imbalanced adapter contributions. This work is the first to empirically reveal the heterogeneous contribution of adapters within Transformer architectures. To address these issues, we propose a selective freezing mechanism: (i) dynamically assessing adapter importance via gradient sensitivity and task-specific contribution; (ii) implementing a staged, progressive freezing strategy; and (iii) incorporating implicit regularization to smooth the loss landscape and improve generalization. Our method maintains or even improves downstream task performance while significantly reducing memory consumption (−42.85%), FLOPs (−34.59%), and training time (−11.82%). The approach achieves superior efficiency without compromising robustness or accuracy, offering a principled and practical solution for resource-constrained adapter tuning.
This work addresses a central challenge in parameter-efficient fine-tuning: identifying the optimal placement of adapters to achieve peak performance with minimal parameters. The authors propose PAGE, a metric based on initial gradient energy analysis, which reveals that adaptation effects are highly concentrated in the down-projection modules of shallow feed-forward networks. Leveraging this insight, they introduce DomLoRA—a method that deploys a single LoRA adapter exclusively in this dominant module. This study is the first to demonstrate the existence, architectural dependency, and task stability of such a dominant adaptation module, establishing a new paradigm for efficient fine-tuning. Experiments show that DomLoRA, using only ~0.7% of the parameters of standard LoRA, consistently outperforms it across diverse tasks—including instruction following, mathematical reasoning, code generation, and multi-turn dialogue—and further enhances the effectiveness of other LoRA variants.
To address parameter redundancy, structural rigidity, and insufficient task adaptability in large language model (LLM) fine-tuning, this paper proposes a learnable-structure adapter tuning method. Under frozen backbone parameters, it jointly optimizes adapter insertion positions, activation paths, and module compositions via differentiable gating and structural sparsity control, dynamically constructing task-specific, efficient substructures. Its key innovation lies in formulating architecture search as a differentiable optimization problem, integrating sensitivity analysis to quantify the impact of sparsity, noise, and data perturbations. Experiments across multiple natural language understanding benchmarks demonstrate that the method significantly outperforms mainstream parameter-efficient fine-tuning approaches—achieving higher accuracy, superior parameter compression ratios, and enhanced robustness against input noise.
This study addresses the routing collapse issue in Mixture-of-Experts (MoE) architectures for time-series foundation models, caused by statistical information stripping during instance normalization. To overcome this, we propose RR-MoA, a causal intervention mechanism grounded in mutual information decomposition that reveals the underlying causes of normalization-induced degradation. Specifically, this work pioneers leveraging raw inputs for routing decisions at the pre-normalization stage, combined with a frozen-backbone adapter mixture technique to enable efficient adaptation across heterogeneous data. Our approach transcends conventional MoE optimization bottlenecks, empirically validates the "freezing paradox," and demonstrates strong cross-backbone generalizability. Extensive evaluation across 54 comparative experiments shows that RR-MoA consistently outperforms both LoRA and full-parameter fine-tuning, achieving state-of-the-art performance without exception.
该研究针对MoE模型参数高效微调中适配器碎片化问题,提出ACE方法,通过整合冗余专家的适配器并执行分组计算,提高准确率和训练速度。
This work investigates the proportion of parameters in neural networks that are truly necessary for encoding task-specific information and proposes a method that trains only extremely low-rank LoRA adapters while keeping the backbone network entirely frozen and randomly initialized. The approach is validated across diverse architectures and tasks, revealing that task-relevant information resides in an exceptionally low-dimensional subspace. This finding implies that randomly initialized backbones are interchangeable and need only be distributed as random seeds. By linking the saturation rank of LoRA to the intrinsic dimensionality of tasks, the method recovers 96%–100% of full fine-tuning performance using merely 0.5%–40% trainable parameters across nine benchmarks, substantially reducing storage and memory overhead.
Existing parameter-efficient fine-tuning methods struggle to recover the multi-scale local geometric information lost during downsampling in 3D point cloud backbone networks, thereby limiting dense prediction performance. This work proposes a local-aware parameter-efficient fine-tuning framework that, while keeping the backbone frozen, introduces for the first time a hierarchical local feature reconstruction mechanism. This mechanism leverages a multi-resolution local feature pyramid, a local-global semantic fusion module, and a dynamic multi-scale prompt generator to restore fine-grained geometric structures, coupled with a lightweight upsampling segmentation head for efficient adaptation. The approach achieves state-of-the-art performance with only 2.71% (for classification) and 7.69% (for dense prediction) trainable parameters, and scales effectively on PointGPT-L with merely 0.36% additional parameters.
This work addresses the challenge of automatically selecting the optimal parameter-efficient fine-tuning (PEFT) adapter during inference in the absence of task labels. The authors propose a training-free, adapter-agnostic dynamic routing framework that selects adapters at inference time by measuring the distance between the input embedding and the centroid of each adapter’s training data in the latent space. This approach establishes the first universal routing mechanism that requires neither access to internal adapter parameters nor any additional training, offering strong scalability and portability across arbitrary PEFT methods. Experimental results demonstrate that on Llama-3.2-1B-Instruct, the method recovers 97.44% of the oracle performance across 23 NLP tasks and achieves an average selection accuracy of 89.7% when scaled to 44 tasks.