parameter-efficient fine-tuning

Designs and evaluates methods for adapting pretrained models by updating only a small subset of parameters or by adding compact modules (e.g., adapters, low-rank updates, prefix/prompt vectors) instead of full-model fine-tuning. This includes building and analyzing techniques that achieve effective task or domain adaptation while preserving accuracy and fidelity, limiting catastrophic forgetting, and reducing storage and compute requirements during training and deployment.

parameter-efficientfine-tuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.55
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$221K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation

Jan 26, 2025
RA
Reza Akbarian Bafghi
🏛️ University of Colorado | Cruise, LLC | University of California, Riverside

To address catastrophic forgetting, degraded out-of-distribution (OOD) generalization, and high computational overhead in large-model domain adaptation, this paper proposes a parameter-efficient fine-tuning method based on selective activation of LoRA modules. Our core innovation is a learnable binary gating function that enables fine-grained, task-aware sparsity in LoRA updates, integrated within the Task Adaptive Parameter Sharing (TAPS) framework and low-rank decomposition. The method updates only ~5% of parameters. Evaluated on CLIP and DINO-ViT, it reduces trainable parameters by over 95% compared to standard LoRA, maintains or improves OOD accuracy, and significantly mitigates forgetting of prior-task knowledge. To our knowledge, this is the first work within the parameter-efficient fine-tuning (PEFT) paradigm to systematically enhance both OOD robustness and long-term knowledge retention.

Continual LearningKnowledge RetentionResource Efficiency

This work addresses catastrophic forgetting in fine-tuning pretrained models, where newly acquired knowledge overwrites previously learned information. To mitigate this issue, the authors propose a function-preserving model expansion approach that mathematically duplicates and scales parameters of selected Transformer submodules during initialization. This technique enables stable training and faithful retention of original model capabilities without altering the initial functionality. By circumventing the traditional trade-off between plasticity and stability, the method achieves performance comparable to full fine-tuning while expanding only a minimal number of layers. Consequently, it fully preserves the model’s original knowledge and substantially reduces computational overhead.

catastrophic forgettingfine-tuningplasticity-stability trade-off

Maintaining Structural Integrity in Parameter Spaces for Parameter Efficient Fine-tuning

May 23, 2024
CS
Chongjie Si
🏛️ Shanghai Jiao Tong University | Shanghai AI Laboratory

To address structural distortion and topological inconsistency in high-dimensional parameter spaces (e.g., 4D tensors) induced by low-rank approximation in parameter-efficient fine-tuning, this paper proposes a structure-preserving low-rank core space modeling method. Unlike conventional low-rank adapters (e.g., LoRA), which are restricted to linear weight matrices, our approach explicitly models and preserves the intrinsic topological structure of the original high-dimensional parameter space—achieving compact and accurate reconstruction of N-dimensional parameter updates via high-order tensor decomposition. Evaluated across CV, NLP, and multimodal benchmarks, the method yields an average accuracy improvement of 1.8% under identical parameter budgets, while reducing structural distortion by 37%, significantly outperforming existing baselines.

Enable parameter-efficient fine-tuning across diverse dimensional spacesModel changes via low-rank core space with consistent topologyPreserve structural integrity in high-dimensional parameter spaces

In large-scale pretraining, learning rate scheduling critically influences both training efficiency and model performance. This work proposes two paradigms—Fitting and Transfer. The Fitting paradigm establishes, for the first time, a scaling law for learning rate search factors, reducing hyperparameter tuning complexity from O(n³) to O(n·C_D·C_η). The Transfer paradigm extends μTransfer to Mixture-of-Experts (MoE) architectures and generalizes it across multiple hyperparameter dimensions, including depth, weight decay, and token length. Empirical results demonstrate that while μTransfer exhibits limited scalability in large-scale settings, the Fitting paradigm—grounded in the derived scaling law—offers superior scalability and practicality, providing a systematic guideline for hyperparameter tuning in industrial-scale pretraining.

hyperparameter optimizationlarge-scale pre-traininglearning rate

An Empirical Analysis of Forgetting in Pre-trained Models with Incremental Low-Rank Updates

May 28, 2024
AS
Albin Soutif-Cormerais
🏛️ Computer Vision Center | Universitat Autònoma de Barcelona | University of Florence

This work investigates how low-rank adaptation (LoRA) parameters influence catastrophic forgetting during fine-tuning. We systematically merge LoRA adapter weights back into the backbone to quantitatively analyze forgetting dynamics across pretraining and downstream tasks, as well as changes in model plasticity. We identify, for the first time, a “contextual forgetting” phenomenon in Vision Transformers (ViTs)—characterized by task-dependent degradation of local features—distinct from the global forgetting observed in ResNets and unreported in prior continual learning literature. Moreover, we reveal that LoRA rank exerts a dual regulatory effect: excessively low ranks exacerbate pretrained knowledge forgetting, while excessively high ranks impair downstream adaptability. Experiments span diverse multi-task continual learning scenarios. Our findings provide theoretical foundations and principled guidelines for rank selection in efficient, sustainable visual model adaptation.

Contextual forgetting in vision transformers vs residual networksEffect of LoRA rank on downstream task plasticityImpact of LoRA rank on pretraining task forgetting

Latest Papers

What's happening recently
View more

Existing theoretical frameworks struggle to explain why larger-scale pre-trained models substantially reduce sample complexity on downstream tasks. This work proposes a novel theoretical framework—termed “caulking”—inspired by parameter-efficient fine-tuning methods such as adapters, low-rank adaptation, and partial fine-tuning. It establishes, for the first time, a provable relationship between the scale of pre-trained models and the sample complexity of downstream tasks. By rigorously linking stronger pre-training capabilities to reduced data requirements in transfer learning, this study not only addresses a critical gap in current theoretical understanding but also provides a solid foundation for empirically observed scaling laws, demonstrating that enhanced pre-training capacity can significantly decrease the number of samples needed for effective downstream adaptation.

downstream taskspre-trained modelssample complexity

Large language models (LLMs) face significant challenges in task adaptation under resource-constrained and closed-source API settings, where conventional parameter-efficient fine-tuning (PEFT) methods are inapplicable due to their reliance on direct model parameter access and high computational overhead. Method: This paper proposes a lightweight, parameter-free knowledge injection framework that enables task-specific adaptation without accessing the LLM’s internal parameters. Its core innovation is the “Specialized Small Model (SSM) Collaboration Paradigm,” integrating knowledge distillation from the LLM, distribution-aware task modeling, and zero-parameter coupling between the SSM and the LLM. Contribution/Results: Experiments demonstrate that our approach matches PEFT-level performance across diverse downstream tasks while reducing GPU memory consumption by over 90% and inference latency by 85%. Crucially, it operates entirely within black-box API environments—requiring no model weights, gradients, or architectural access—thus enabling seamless integration with proprietary, closed-source LLM APIs.

Adapts large models to tasks without accessing their parameters.Enhances performance on specific distributions using small models.Reduces resource costs for fine-tuning in constrained environments.

This study addresses the susceptibility of low-rank parameter-efficient fine-tuning (PEFT) methods to catastrophic forgetting in continual learning, a phenomenon whose underlying mechanisms remain poorly understood. Through empirical analysis of the geometric structure and parametrization of update subspaces in representative low-rank and tensor decomposition approaches—including LoRA, LoRETTA, and WeGeFT—the work identifies subspace design as a critical factor governing forgetting behavior. The findings reveal that methods employing tensor decomposition (e.g., LoRETTA) or structurally aligned parametrization (e.g., WeGeFT) substantially mitigate catastrophic forgetting even under extremely limited parameter budgets, outperforming conventional shared-subspace strategies. These insights provide both theoretical grounding and practical guidance for designing efficient adaptation mechanisms in continual learning scenarios.

catastrophic forgettingcontinual learninglow-rank decomposition

This work addresses catastrophic forgetting in continual learning with pre-trained models when access to previous task data is prohibited. The authors propose a structured low-rank adaptation method grounded in geometric redundancy of pre-trained weights. By analyzing the intrinsic geometric structure of the pre-trained weight space, they identify a protected subspace for parameter updates and formulate the update as \( \Delta W = BAQ^\top \), where frozen matrices \( B \) and \( Q \) project trainable low-rank matrix \( A \) exclusively onto redundant directions. This approach is the first to leverage geometric redundancy to explicitly locate plasticity regions, enabling a controllable trade-off between plasticity and stability without requiring data replay. Experimental results demonstrate that the method effectively suppresses functional drift and significantly improves retention of performance on prior tasks, even under worst-case scenarios.

continual learningdata-free adaptationfoundation models

Mitigating Forgetting in Low Rank Adaptation

Dec 19, 2025
JS
Joanna Sliwa
🏛️ University of Tübingen | University of Cambridge

To address catastrophic forgetting in LoRA-based fine-tuning, this paper proposes LaLoRA—the first method to introduce the Laplace approximation into the LoRA weight space, enabling lightweight weight-space regularization. By constraining parameter updates along high-curvature directions, LaLoRA preserves pretraining knowledge while enhancing downstream task performance. Crucially, regularization operates solely on the low-rank incremental matrices, incurring no inference overhead and supporting tunable trade-offs between learning and forgetting. Its core innovation lies in modeling LoRA parameter confidence via loss curvature estimation, unifying parameter efficiency, knowledge stability, and robustness. In mathematical reasoning fine-tuning experiments on Llama, LaLoRA significantly improves the forgetting–performance trade-off: regularization strength directly controls forgetting extent, and the method demonstrates strong robustness to data sampling variations and hyperparameter choices.

Controls learning-forgetting trade-off via regularization strengthMitigates catastrophic forgetting in LoRA fine-tuningPreserves prior knowledge while enabling efficient learning

Hot Scholars

YL

Yuhang Liu

The University of Adelaide
Representation LearningLLMsLatent Variable ModelsResponsible AI
YS

Yiren Song

PH.D student, National University of Singapore
Generative AIDiffusionUnified model
MY

Mohammad Yaqub

Researcher in Biomedical Engineering, Associate professor at MBZUAI
Artificial IntelligenceMedical Image AnalysisMachine LearningDeep learning
XB

Xiang Bai

Huazhong University of Science and Technology (HUST)
Computer VisionOCR
AZ

An Zhang

University of Science and Technology
Generative ModelsTrustworthy AIAgentic AIRecommender System