parameter-efficient transformer tuning

Designs and implements methods to adapt pretrained transformer models while training far fewer parameters than full fine‑tuning, by injecting compact modules or updates (e.g., low‑rank adapters, additive or sparse parameter updates, prompt vectors) that leave the base transformer weights largely unchanged. Builds and evaluates these parameter‑efficient modules and analyzes tradeoffs among number of trainable parameters, predictive accuracy, compute and storage requirements when integrated into encoder/decoder architectures.

parameter-efficienttransformertuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.68
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

PARA: Parameter-Efficient Fine-tuning with Prompt-Aware Representation Adjustment

Feb 03, 2025
ZL
Zequan Liu
🏛️ RWTH Aachen University | University of Pennsylvania | Southern University of Science and Technology | University of Hong Kong | Carnegie Mellon University

To address the challenge of balancing parameter efficiency and cross-task generalization in single-backbone multi-task fine-tuning, this paper proposes PARA, a lightweight parameter-efficient fine-tuning (PEFT) method. Its core innovation lies in introducing a prompt-aware vector generator into each Transformer layer—enabling fine-grained, low-overhead dynamic adaptation of hidden-layer representations. PARA further integrates intra-layer feature reweighting with strategic parameter freezing, achieving substantial gains in multi-task generalization while adding only 0.1% extra parameters—comparable to LoRA. Empirical evaluation on comprehensive multi-task benchmarks demonstrates that PARA consistently outperforms existing PEFT methods. Moreover, under single-backbone multi-tenant deployment, PARA accelerates inference by 17% relative to LoRA, confirming its superior efficiency–effectiveness trade-off.

Multi-task LearningParameter EfficiencyPrompt Tuning

Module-Aware Parameter-Efficient Machine Unlearning on Transformers

Aug 24, 2025
WB
Wenjie Bao
🏛️ Zhejiang University | UNC Greensboro

Existing parameter-efficient unlearning methods neglect the structural characteristics of Transformer modules, hindering precise identification of critical parameters and thus limiting unlearning performance. To address this, we propose MAPE-Unlearn, a module-aware parameter-efficient unlearning framework. It introduces, for the first time, a module-aware mechanism that employs learnable dual masks—separately applied to attention heads and feed-forward filters—optimized via warm-started greedy search to accurately locate and freeze task-critical parameters. The objective function is explicitly driven by the unlearning target. Extensive experiments across multiple architectures (BERT, RoBERTa, ViT) and benchmarks (SST-2, AG News, ImageNet) demonstrate that MAPE-Unlearn significantly outperforms state-of-the-art parameter-efficient unlearning methods, achieving superior unlearning accuracy, enhanced model robustness, and strong generalization capability.

Identifying influence-critical parameters in Transformers for unlearningImproving unlearning performance with module-aware parameter efficiencyRemoving specific data influences to comply with privacy regulations

Not All Adapters Matter: Selective Adapter Freezing for Memory-Efficient Fine-Tuning of Language Models

Nov 26, 2024
HS
Hyegang Son
🏛️ Korea University | Arizona State University | Soongsil University

Existing adapter-based fine-tuning methods, though parameter-efficient, suffer from substantial memory and computational overhead, prolonged training time, and imbalanced adapter contributions. This work is the first to empirically reveal the heterogeneous contribution of adapters within Transformer architectures. To address these issues, we propose a selective freezing mechanism: (i) dynamically assessing adapter importance via gradient sensitivity and task-specific contribution; (ii) implementing a staged, progressive freezing strategy; and (iii) incorporating implicit regularization to smooth the loss landscape and improve generalization. Our method maintains or even improves downstream task performance while significantly reducing memory consumption (−42.85%), FLOPs (−34.59%), and training time (−11.82%). The approach achieves superior efficiency without compromising robustness or accuracy, offering a principled and practical solution for resource-constrained adapter tuning.

Improve generalization by smoothing loss landscapeMaintain performance while optimizing memory and computationSelectively freeze adapters to reduce resource usage

Computational Limits of Low-Rank Adaptation (LoRA) for Transformer-Based Models

Jun 05, 2024
JY
Jerry Yao-Chieh Hu
🏛️ Northwestern University | USTC | National Center for Theoretical Sciences | Adobe Research

This work addresses the high computational complexity of gradient computation in Low-Rank Adaptation (LoRA) fine-tuning. Under the Strong Exponential Time Hypothesis (SETH), we establish the first computational phase-transition theory for LoRA gradient computation: we prove that a subquadratic approximation algorithm exists when the LoRA rank satisfies a norm-sharing upper-bound threshold; further, we design a chain-based low-rank approximation scheme enabling near-linear gradient updates. Methodologically, we propose a term-wise controlled analytical framework integrating fine-grained complexity analysis, hierarchical low-rank gradient approximation, and selective adaptation of attention weight subsets (Q/V/K). Our core contributions are: (1) the first characterization of the efficiency phase-transition threshold for LoRA gradient computation; (2) the identification of sufficient conditions for near-linear solvability; and (3) the first theory-driven acceleration pathway for low-rank adaptive methods.

Analyze computational limits of LoRA for transformer fine-tuningIdentify phase transition in LoRA efficiency under SETHProve existence of linear algorithms via low-rank gradients

Mini-Ensemble Low-Rank Adapters for Parameter-Efficient Fine-Tuning

Feb 27, 2024
PR
Pengjie Ren
🏛️ Shandong University | Leiden University | University of Amsterdam | Centrum Wiskunde & Informatica

To address LoRA’s limited generalization under low-rank constraints and its inferiority to full-parameter fine-tuning, this paper proposes MELoRA—a miniature ensemble-based low-rank adaptation method. MELoRA freezes the pretrained LLM weights and trains multiple ultra-lightweight (rank-1) mini-LoRA adapters in parallel, forming an ensemble. By synergistically combining low-rank matrix decomposition with ensemble learning, it significantly enhances representational capacity while drastically reducing trainable parameters. Theoretical analysis establishes a tighter error upper bound for MELoRA compared to standard LoRA. Empirical evaluation across diverse tasks shows that MELoRA outperforms LoRA on natural language understanding with only 1/8 of its parameters, and achieves comparable or superior performance on instruction-following tasks using merely 1/36 of LoRA’s parameters. To our knowledge, MELoRA is the first PEFT framework to integrate ensemble learning into low-rank adaptation architectures, effectively overcoming the generalization bottleneck inherent in single-adapter approaches.

Enhances diversity among mini-adapters for better task adaptationImproves generalization in parameter-efficient fine-tuning with low-rank adaptersReduces trainable parameters while maintaining high model performance

Latest Papers

What's happening recently
View more

This work addresses the limitation of fixed-rank constraints in parameter-efficient fine-tuning, which fail to accommodate the heterogeneous rank requirements across different layers of neural networks. The authors propose LR-LoRA, a novel approach that introduces a learnable rank mechanism within the LoRA framework, enabling differentiable and dynamic optimization of the rank for each adapter layer. This method reveals a systematic disparity in rank demands between attention and MLP layers in Transformers, thereby providing a more flexible and effective inductive bias. Experimental results demonstrate that LR-LoRA significantly outperforms existing parameter-efficient fine-tuning methods across multiple benchmarks for language understanding and commonsense reasoning, achieving state-of-the-art performance.

adapter rankinductive biaslow-rank adaptation

Existing parameter-efficient fine-tuning (PEFT) methods struggle to simultaneously achieve high-rank weight updates and vector-level parameter efficiency. To address this, we propose TeRA: a high-rank adapter based on randomized tensor networks. TeRA decouples the rank of the adaptation matrix from its trainable parameter count by freezing shared, randomly tensorized factors and learning only layer-specific diagonal scaling vectors. It employs a Tucker-like decomposition, dynamically generating high-rank adaptation matrices via tensor products—enabling high-rank updates while maintaining vector-level parameter counts comparable to LoRA or IA3. Experiments demonstrate that TeRA consistently outperforms existing high-rank adapters across multiple benchmarks. Theoretical analysis and ablation studies further validate its design efficacy. This work is the first to introduce tensor networks into PEFT, breaking the fundamental trade-off between representational capacity and parameter efficiency.

Achieving high-rank adaptation without sacrificing parameter efficiencyEnabling high-rank weight updates with vector-based parameter efficiencyResolving trade-off between expressivity and parameter reduction in PEFT

PreLoRA: Hybrid Pre-training of Vision Transformers with Full Training and Low-Rank Adapters

Sep 25, 2025
KK
Krishu K Thapa
🏛️ Washington State University | Argonne National Laboratory

To address the high computational resource demands of training large-scale Vision Transformers (ViTs), this paper proposes Dynamic Low-Rank Adaptation (Dynamic LoRA). Unlike static approaches, Dynamic LoRA initially performs full-parameter fine-tuning, then dynamically identifies layer-wise convergence by monitoring weight update magnitudes and adaptively transitions to hierarchical low-rank adaptation. Its key innovations include a module-level rank allocation strategy and a hyperparameter-driven phase-switching mechanism. Evaluated on ViT-Large, the method achieves lossless accuracy relative to full fine-tuning. Experiments demonstrate that Dynamic LoRA reduces trainable parameters by 90% (to 10% of full fine-tuning), triples training throughput, shortens per-epoch training time by 1.5×, and decreases GPU memory consumption by 20%. It consistently outperforms both static LoRA and full fine-tuning across efficiency and accuracy metrics.

Dynamically switches from full training to low-rank adaptationMaintains accuracy while cutting parameters and improving efficiencyReduces resource-intensive training of large vision transformers

This study addresses the unclear performance gap between dense and sparse Mixture-of-Experts (MoE) models at extremely small scales (<25M parameters). Under a unified LLaMA-style pretraining setup, the authors systematically compare both architectures while fixing all hyperparameters except model width, aligning either active or total parameter counts. To stabilize sparse training, they incorporate Switch-style load balancing and router z-loss. Results show that when active parameters are matched, MoE achieves lower validation loss than its dense counterpart (1.5788 vs. 1.6545). However, under total parameter alignment—reflecting equal storage footprint—the dense model slightly outperforms MoE (1.5608 vs. 1.5788), revealing that MoE struggles to surpass dense models of equivalent capacity at very small scales.

dense vs sparsemixture-of-expertsparameter matching

Hot Scholars

MM

Michael Muehlebach

Max Planck Institute for Intelligent Systems
Machine LearningOptimizationDynamical Systems
CV

Claire Vernade

University of Tuebingen
Machine LearningMulti-armed banditreinforcement learning