Adapt Once, Thrive with Updates: Transferable Parameter-Efficient Fine-Tuning on Evolving Base Models

📅 2025-06-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Foundation model updates cause severe performance degradation in existing parameter-efficient fine-tuning (PEFT) modules, necessitating costly retraining. Method: This work first identifies that feed-forward networks (FFNs) encode update-sensitive knowledge, while attention mechanisms retain stable task-specific patterns. Leveraging this insight, we propose a transferable PEFT architecture that decouples task patterns from base-model knowledge. Our approach comprises attention stability analysis, FFN knowledge sensitivity modeling, parameter-decoupled structural design, and theoretical convergence guarantees. Results: Evaluated across seven foundation models and twelve datasets, our method enables zero-cost migration of legacy PEFT modules to updated models, achieving >98% average performance retention—significantly reducing operational overhead during model iteration. The core contribution is the first retraining-free PEFT transfer paradigm explicitly designed for foundation model version evolution.

Technology Category

Machine Learning: Deep Neural Architectures and Foundation ModelsComputer Vision: Large Vision ModelsNatural Language Processing: (Large) Language Models

Application Category

Graph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Parameter-efficient fine-tuning (PEFT) has become a common method for fine-tuning large language models, where a base model can serve multiple users through PEFT module switching. To enhance user experience, base models require periodic updates. However, once updated, PEFT modules fine-tuned on previous versions often suffer substantial performance degradation on newer versions. Re-tuning these numerous modules to restore performance would incur significant computational costs. Through a comprehensive analysis of the changes that occur during base model updates, we uncover an interesting phenomenon: continual training primarily affects task-specific knowledge stored in Feed-Forward Networks (FFN), while having less impact on the task-specific pattern in the Attention mechanism. Based on these findings, we introduce Trans-PEFT, a novel approach that enhances the PEFT module by focusing on the task-specific pattern while reducing its dependence on certain knowledge in the base model. Further theoretical analysis supports our approach. Extensive experiments across 7 base models and 12 datasets demonstrate that Trans-PEFT trained modules can maintain performance on updated base models without re-tuning, significantly reducing maintenance overhead in real-world applications.
Problem

Research questions and friction points this paper is trying to address.

PEFT modules degrade after base model updates
Re-tuning PEFT modules is computationally expensive
Trans-PEFT maintains performance without re-tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Trans-PEFT focuses on task-specific patterns
Reduces dependency on base model knowledge
Maintains performance without re-tuning modules
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
N
Naibin Gu
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
P
Peng Fu
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
X
Xiyu Liu
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
K
Ke Ma
School of Electronic, Electrical and Communication Engineering, UCAS, Beijing, China
Z
Zheng Lin
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China
Weiping Wang
Weiping Wang
School of Information Science and Engineering, Central South University
Computer NetworkNetwork Security