side-stem adapter tuning

Designs and implements lightweight adapter modules added as auxiliary side stems or input paths to a frozen pretrained backbone, along with the tuning procedures that update only those adapter parameters (often low-rank) to adapt the model to a new task. This includes building the side-stem architecture, attention or gating mechanisms for combining stem and backbone activations, and measuring parameter-efficiency versus task performance.

side-stemadaptertuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.58
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Structural Priors and Modular Adapters in the Composable Fine-Tuning Algorithm of Large-Scale Models

Nov 06, 2025
YW
Yuxiao Wang
🏛️ University of Pennsylvania | Washington University in St. Louis | Stevens Institute of Technology | University of Southern California

Large-scale pretrained models face high computational overhead and structural instability during multi-task adaptation. Method: This paper proposes a composable fine-tuning framework that integrates graph-structured task priors with modular adapters. It constructs a task-relation graph to model inter-task dependencies, leveraging this structured prior to guide low-rank adapter parameter allocation and dynamic routing. The framework incorporates plug-and-play adapter design, relation-matrix regularization, and temperature- and gating-based control mechanisms to mitigate path conflicts and redundant computation. Contributions/Results: Experiments demonstrate significant improvements in task prediction accuracy and adapter assignment precision. The method exhibits strong robustness under hyperparameter, environmental, and data perturbations, achieving both high performance and parameter efficiency. It establishes a new paradigm for multi-task adaptation—characterized by interpretability, reusability, and structural stability—without compromising scalability or practicality.

Addressing structural instability through graph-based dependency modelingImproving parameter efficiency and training stability via modular adaptersReducing computational costs in multi-task adaptation of large-scale models

MoRe Fine-Tuning with 10x Fewer Parameters

Aug 30, 2024
WT
Wenxuan Tan
🏛️ University of Wisconsin—Madison

Existing parameter-efficient fine-tuning (PEFT) methods—such as LoRA—rely on heuristic adapter architectures, suffering from poor generalization and limited transferability across models. Method: We propose the first learnable rectangular adapter search framework grounded in Monarch matrices—the first application of Monarch structure to PEFT—supported by theoretical analysis demonstrating superior expressivity over LoRA. Our approach employs differentiable neural architecture search to automatically discover optimal lightweight adapter topologies, eliminating manual specification of rank or module shape, and integrates low-parameter adapter design with efficient fine-tuning strategies. Contribution/Results: On multi-task and multi-model benchmarks, our method significantly outperforms state-of-the-art PEFT approaches, achieving comparable or superior performance using only 5% of LoRA’s parameters. It delivers both strong cross-task/model generalization and exceptional parameter efficiency.

Enhancing performance with fewer parameters in PEFT techniquesOptimizing adapter architectures for parameter-efficient fine-tuningReducing reliance on heuristics in low-rank adapters (LoRA)

Not All Adapters Matter: Selective Adapter Freezing for Memory-Efficient Fine-Tuning of Language Models

Nov 26, 2024
HS
Hyegang Son
🏛️ Korea University | Arizona State University | Soongsil University

Existing adapter-based fine-tuning methods, though parameter-efficient, suffer from substantial memory and computational overhead, prolonged training time, and imbalanced adapter contributions. This work is the first to empirically reveal the heterogeneous contribution of adapters within Transformer architectures. To address these issues, we propose a selective freezing mechanism: (i) dynamically assessing adapter importance via gradient sensitivity and task-specific contribution; (ii) implementing a staged, progressive freezing strategy; and (iii) incorporating implicit regularization to smooth the loss landscape and improve generalization. Our method maintains or even improves downstream task performance while significantly reducing memory consumption (−42.85%), FLOPs (−34.59%), and training time (−11.82%). The approach achieves superior efficiency without compromising robustness or accuracy, offering a principled and practical solution for resource-constrained adapter tuning.

Improve generalization by smoothing loss landscapeMaintain performance while optimizing memory and computationSelectively freeze adapters to reduce resource usage

This work addresses a central challenge in parameter-efficient fine-tuning: identifying the optimal placement of adapters to achieve peak performance with minimal parameters. The authors propose PAGE, a metric based on initial gradient energy analysis, which reveals that adaptation effects are highly concentrated in the down-projection modules of shallow feed-forward networks. Leveraging this insight, they introduce DomLoRA—a method that deploys a single LoRA adapter exclusively in this dominant module. This study is the first to demonstrate the existence, architectural dependency, and task stability of such a dominant adaptation module, establishing a new paradigm for efficient fine-tuning. Experiments show that DomLoRA, using only ~0.7% of the parameters of standard LoRA, consistently outperforms it across diverse tasks—including instruction following, mathematical reasoning, code generation, and multi-turn dialogue—and further enhances the effectiveness of other LoRA variants.

adapter placementdominant adaptation modulegradient energy

Structure-Learnable Adapter Fine-Tuning for Parameter-Efficient Large Language Models

Sep 03, 2025
MG
Ming Gong
🏛️ University of Pennsylvania | Georgia Institute of Technology | Independent Author | University of California, Berkeley

To address parameter redundancy, structural rigidity, and insufficient task adaptability in large language model (LLM) fine-tuning, this paper proposes a learnable-structure adapter tuning method. Under frozen backbone parameters, it jointly optimizes adapter insertion positions, activation paths, and module compositions via differentiable gating and structural sparsity control, dynamically constructing task-specific, efficient substructures. Its key innovation lies in formulating architecture search as a differentiable optimization problem, integrating sensitivity analysis to quantify the impact of sparsity, noise, and data perturbations. Experiments across multiple natural language understanding benchmarks demonstrate that the method significantly outperforms mainstream parameter-efficient fine-tuning approaches—achieving higher accuracy, superior parameter compression ratios, and enhanced robustness against input noise.

Enhances task adaptability through learnable adapter structuresImproves robustness against noise and data perturbationsReduces parameter redundancy in fine-tuning large models

Latest Papers

What's happening recently
View more

This study addresses the routing collapse issue in Mixture-of-Experts (MoE) architectures for time-series foundation models, caused by statistical information stripping during instance normalization. To overcome this, we propose RR-MoA, a causal intervention mechanism grounded in mutual information decomposition that reveals the underlying causes of normalization-induced degradation. Specifically, this work pioneers leveraging raw inputs for routing decisions at the pre-normalization stage, combined with a frozen-backbone adapter mixture technique to enable efficient adaptation across heterogeneous data. Our approach transcends conventional MoE optimization bottlenecks, empirically validates the "freezing paradox," and demonstrates strong cross-backbone generalizability. Extensive evaluation across 54 comparative experiments shows that RR-MoA consistently outperforms both LoRA and full-parameter fine-tuning, achieving state-of-the-art performance without exception.

AdapterInstance NormalizationMixture of Experts

This work investigates the proportion of parameters in neural networks that are truly necessary for encoding task-specific information and proposes a method that trains only extremely low-rank LoRA adapters while keeping the backbone network entirely frozen and randomly initialized. The approach is validated across diverse architectures and tasks, revealing that task-relevant information resides in an exceptionally low-dimensional subspace. This finding implies that randomly initialized backbones are interchangeable and need only be distributed as random seeds. By linking the saturation rank of LoRA to the intrinsic dimensionality of tasks, the method recovers 96%–100% of full fine-tuning performance using merely 0.5%–40% trainable parameters across nine benchmarks, substantially reducing storage and memory overhead.

intrinsic dimensionalitylow-rank adaptationneural network parameters

Existing parameter-efficient fine-tuning methods struggle to recover the multi-scale local geometric information lost during downsampling in 3D point cloud backbone networks, thereby limiting dense prediction performance. This work proposes a local-aware parameter-efficient fine-tuning framework that, while keeping the backbone frozen, introduces for the first time a hierarchical local feature reconstruction mechanism. This mechanism leverages a multi-resolution local feature pyramid, a local-global semantic fusion module, and a dynamic multi-scale prompt generator to restore fine-grained geometric structures, coupled with a lightweight upsampling segmentation head for efficient adaptation. The approach achieves state-of-the-art performance with only 2.71% (for classification) and 7.69% (for dense prediction) trainable parameters, and scales effectively on PointGPT-L with merely 0.36% additional parameters.

3D representation learninglocal geometry recoverymulti-scale locality

This work addresses the challenge of automatically selecting the optimal parameter-efficient fine-tuning (PEFT) adapter during inference in the absence of task labels. The authors propose a training-free, adapter-agnostic dynamic routing framework that selects adapters at inference time by measuring the distance between the input embedding and the centroid of each adapter’s training data in the latent space. This approach establishes the first universal routing mechanism that requires neither access to internal adapter parameters nor any additional training, offering strong scalability and portability across arbitrary PEFT methods. Experimental results demonstrate that on Llama-3.2-1B-Instruct, the method recovers 97.44% of the oracle performance across 23 NLP tasks and achieves an average selection accuracy of 89.7% when scaled to 44 tasks.

adapter selectioninference-time routingmodel ecosystem