Prefix-Tuning+: Modernizing Prefix-Tuning through Attention Independent Prefix Data

πŸ“… 2025-06-16
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Prefix-Tuning suffers significant performance degradation on modern large language models (LLMs), primarily due to intrinsic competition between the input sequence and learnable prefixes within attention headsβ€”leading to suppression of prefix importance. Method: This work first identifies this mechanism and proposes Attention-agnostic Prefix (AAP), a novel architecture that decouples prefix modules from attention computation entirely, recasting them as context-aware prefix generators independent of attention heads. AAP comprises three key components: (i) Transformer-based prefix injection reconstruction, (ii) attention-head decoupling design, and (iii) dynamic prefix construction strategy. Contribution/Results: Across diverse benchmark tasks, AAP consistently outperforms standard Prefix-Tuning and achieves generalization capability on par with LoRA. Empirical results validate AAP as an effective, broadly applicable paradigm for parameter-efficient fine-tuning (PEFT), establishing its potential as a next-generation mainstream PEFT framework.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsHumans and AI: Human-Aware Planning and Behavior Prediction

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingWeb Mining and Content Analysis: Large pretrained models with web data
πŸ“ Abstract
Parameter-Efficient Fine-Tuning (PEFT) methods have become crucial for rapidly adapting large language models (LLMs) to downstream tasks. Prefix-Tuning, an early and effective PEFT technique, demonstrated the ability to achieve performance comparable to full fine-tuning with significantly reduced computational and memory overhead. However, despite its earlier success, its effectiveness in training modern state-of-the-art LLMs has been very limited. In this work, we demonstrate empirically that Prefix-Tuning underperforms on LLMs because of an inherent tradeoff between input and prefix significance within the attention head. This motivates us to introduce Prefix-Tuning+, a novel architecture that generalizes the principles of Prefix-Tuning while addressing its shortcomings by shifting the prefix module out of the attention head itself. We further provide an overview of our construction process to guide future users when constructing their own context-based methods. Our experiments show that, across a diverse set of benchmarks, Prefix-Tuning+ consistently outperforms existing Prefix-Tuning methods. Notably, it achieves performance on par with the widely adopted LoRA method on several general benchmarks, highlighting the potential modern extension of Prefix-Tuning approaches. Our findings suggest that by overcoming its inherent limitations, Prefix-Tuning can remain a competitive and relevant research direction in the landscape of parameter-efficient LLM adaptation.
Problem

Research questions and friction points this paper is trying to address.

Improving Prefix-Tuning for modern LLMs
Addressing attention head tradeoff in Prefix-Tuning
Enhancing parameter-efficient fine-tuning performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Shifts prefix module out of attention head
Generalizes principles of Prefix-Tuning
Outperforms existing Prefix-Tuning methods
H
Haonan Wang
National University of Singapore
Brian Chen
Brian Chen
Google DeepMind; Samsung Research America; Columbia University
Computer visionVision and LanguageMultimodal Learning
S
Siquan Li
National University of Singapore
L
Liang Xinhe
National University of Singapore
Tianyang Hu
Tianyang Hu
Assistant Professor, The Chinese University of Hong Kong, Shenzhen
Deep LearningMaching LearningStatistics
H
Hwee Kuan Lee
National University of Singapore, Bioinformatics Institute, A*STAR, Nanyang Technological University, Singapore Eye Research Institute, Singapore International Research Laboratory on AI, and Singapore Institute for Clinical Sciences
Kenji Kawaguchi
Kenji Kawaguchi
Presidential Young Professor, National University of Singapore
LLMsLarge language modelDeep learningAI