Score
Designs and implements methods to adapt pretrained transformer models while training far fewer parameters than full fine‑tuning, by injecting compact modules or updates (e.g., low‑rank adapters, additive or sparse parameter updates, prompt vectors) that leave the base transformer weights largely unchanged. Builds and evaluates these parameter‑efficient modules and analyzes tradeoffs among number of trainable parameters, predictive accuracy, compute and storage requirements when integrated into encoder/decoder architectures.
Large language models (LLMs) face significant challenges in full-parameter fine-tuning under constrained GPU memory and computational resources, hindering efficient adaptation to downstream tasks. To address this, this work systematically surveys parameter-efficient fine-tuning (PEFT) methodologies and proposes the first unified conceptual framework—comprehensively covering theoretical foundations, algorithmic taxonomies (e.g., LoRA, Adapter, Prompt/Prefix Tuning), cross-modal extensions, and emerging trends. Distinct from fragmented surveys, our framework explicitly articulates theoretical interconnections and practical applicability boundaries across methods, unifying representative paradigms from both NLP and multimodal learning. We further release an open-source, structured knowledge graph encoding these insights. The resulting framework substantially lowers the barrier to lightweight LLM adaptation, offering researchers and practitioners a reusable, transferable technical guide. By bridging theoretical analysis with engineering pragmatism, this work accelerates the transition of PEFT from methodological exploration to scalable, production-ready deployment.
To address the challenge of balancing parameter efficiency and cross-task generalization in single-backbone multi-task fine-tuning, this paper proposes PARA, a lightweight parameter-efficient fine-tuning (PEFT) method. Its core innovation lies in introducing a prompt-aware vector generator into each Transformer layer—enabling fine-grained, low-overhead dynamic adaptation of hidden-layer representations. PARA further integrates intra-layer feature reweighting with strategic parameter freezing, achieving substantial gains in multi-task generalization while adding only 0.1% extra parameters—comparable to LoRA. Empirical evaluation on comprehensive multi-task benchmarks demonstrates that PARA consistently outperforms existing PEFT methods. Moreover, under single-backbone multi-tenant deployment, PARA accelerates inference by 17% relative to LoRA, confirming its superior efficiency–effectiveness trade-off.
Existing parameter-efficient unlearning methods neglect the structural characteristics of Transformer modules, hindering precise identification of critical parameters and thus limiting unlearning performance. To address this, we propose MAPE-Unlearn, a module-aware parameter-efficient unlearning framework. It introduces, for the first time, a module-aware mechanism that employs learnable dual masks—separately applied to attention heads and feed-forward filters—optimized via warm-started greedy search to accurately locate and freeze task-critical parameters. The objective function is explicitly driven by the unlearning target. Extensive experiments across multiple architectures (BERT, RoBERTa, ViT) and benchmarks (SST-2, AG News, ImageNet) demonstrate that MAPE-Unlearn significantly outperforms state-of-the-art parameter-efficient unlearning methods, achieving superior unlearning accuracy, enhanced model robustness, and strong generalization capability.
Existing adapter-based fine-tuning methods, though parameter-efficient, suffer from substantial memory and computational overhead, prolonged training time, and imbalanced adapter contributions. This work is the first to empirically reveal the heterogeneous contribution of adapters within Transformer architectures. To address these issues, we propose a selective freezing mechanism: (i) dynamically assessing adapter importance via gradient sensitivity and task-specific contribution; (ii) implementing a staged, progressive freezing strategy; and (iii) incorporating implicit regularization to smooth the loss landscape and improve generalization. Our method maintains or even improves downstream task performance while significantly reducing memory consumption (−42.85%), FLOPs (−34.59%), and training time (−11.82%). The approach achieves superior efficiency without compromising robustness or accuracy, offering a principled and practical solution for resource-constrained adapter tuning.
This work addresses the high computational complexity of gradient computation in Low-Rank Adaptation (LoRA) fine-tuning. Under the Strong Exponential Time Hypothesis (SETH), we establish the first computational phase-transition theory for LoRA gradient computation: we prove that a subquadratic approximation algorithm exists when the LoRA rank satisfies a norm-sharing upper-bound threshold; further, we design a chain-based low-rank approximation scheme enabling near-linear gradient updates. Methodologically, we propose a term-wise controlled analytical framework integrating fine-grained complexity analysis, hierarchical low-rank gradient approximation, and selective adaptation of attention weight subsets (Q/V/K). Our core contributions are: (1) the first characterization of the efficiency phase-transition threshold for LoRA gradient computation; (2) the identification of sufficient conditions for near-linear solvability; and (3) the first theory-driven acceleration pathway for low-rank adaptive methods.
To address LoRA’s limited generalization under low-rank constraints and its inferiority to full-parameter fine-tuning, this paper proposes MELoRA—a miniature ensemble-based low-rank adaptation method. MELoRA freezes the pretrained LLM weights and trains multiple ultra-lightweight (rank-1) mini-LoRA adapters in parallel, forming an ensemble. By synergistically combining low-rank matrix decomposition with ensemble learning, it significantly enhances representational capacity while drastically reducing trainable parameters. Theoretical analysis establishes a tighter error upper bound for MELoRA compared to standard LoRA. Empirical evaluation across diverse tasks shows that MELoRA outperforms LoRA on natural language understanding with only 1/8 of its parameters, and achieves comparable or superior performance on instruction-following tasks using merely 1/36 of LoRA’s parameters. To our knowledge, MELoRA is the first PEFT framework to integrate ensemble learning into low-rank adaptation architectures, effectively overcoming the generalization bottleneck inherent in single-adapter approaches.
This work addresses the limitation of fixed-rank constraints in parameter-efficient fine-tuning, which fail to accommodate the heterogeneous rank requirements across different layers of neural networks. The authors propose LR-LoRA, a novel approach that introduces a learnable rank mechanism within the LoRA framework, enabling differentiable and dynamic optimization of the rank for each adapter layer. This method reveals a systematic disparity in rank demands between attention and MLP layers in Transformers, thereby providing a more flexible and effective inductive bias. Experimental results demonstrate that LR-LoRA significantly outperforms existing parameter-efficient fine-tuning methods across multiple benchmarks for language understanding and commonsense reasoning, achieving state-of-the-art performance.
Existing parameter-efficient fine-tuning (PEFT) methods struggle to simultaneously achieve high-rank weight updates and vector-level parameter efficiency. To address this, we propose TeRA: a high-rank adapter based on randomized tensor networks. TeRA decouples the rank of the adaptation matrix from its trainable parameter count by freezing shared, randomly tensorized factors and learning only layer-specific diagonal scaling vectors. It employs a Tucker-like decomposition, dynamically generating high-rank adaptation matrices via tensor products—enabling high-rank updates while maintaining vector-level parameter counts comparable to LoRA or IA3. Experiments demonstrate that TeRA consistently outperforms existing high-rank adapters across multiple benchmarks. Theoretical analysis and ablation studies further validate its design efficacy. This work is the first to introduce tensor networks into PEFT, breaking the fundamental trade-off between representational capacity and parameter efficiency.
To address the high computational resource demands of training large-scale Vision Transformers (ViTs), this paper proposes Dynamic Low-Rank Adaptation (Dynamic LoRA). Unlike static approaches, Dynamic LoRA initially performs full-parameter fine-tuning, then dynamically identifies layer-wise convergence by monitoring weight update magnitudes and adaptively transitions to hierarchical low-rank adaptation. Its key innovations include a module-level rank allocation strategy and a hyperparameter-driven phase-switching mechanism. Evaluated on ViT-Large, the method achieves lossless accuracy relative to full fine-tuning. Experiments demonstrate that Dynamic LoRA reduces trainable parameters by 90% (to 10% of full fine-tuning), triples training throughput, shortens per-epoch training time by 1.5×, and decreases GPU memory consumption by 20%. It consistently outperforms both static LoRA and full fine-tuning across efficiency and accuracy metrics.
This study addresses the unclear performance gap between dense and sparse Mixture-of-Experts (MoE) models at extremely small scales (<25M parameters). Under a unified LLaMA-style pretraining setup, the authors systematically compare both architectures while fixing all hyperparameters except model width, aligning either active or total parameter counts. To stabilize sparse training, they incorporate Switch-style load balancing and router z-loss. Results show that when active parameters are matched, MoE achieves lower validation loss than its dense counterpart (1.5788 vs. 1.6545). However, under total parameter alignment—reflecting equal storage footprint—the dense model slightly outperforms MoE (1.5608 vs. 1.5788), revealing that MoE struggles to surpass dense models of equivalent capacity at very small scales.