π€ AI Summary
Prefix-Tuning suffers significant performance degradation on modern large language models (LLMs), primarily due to intrinsic competition between the input sequence and learnable prefixes within attention headsβleading to suppression of prefix importance.
Method: This work first identifies this mechanism and proposes Attention-agnostic Prefix (AAP), a novel architecture that decouples prefix modules from attention computation entirely, recasting them as context-aware prefix generators independent of attention heads. AAP comprises three key components: (i) Transformer-based prefix injection reconstruction, (ii) attention-head decoupling design, and (iii) dynamic prefix construction strategy.
Contribution/Results: Across diverse benchmark tasks, AAP consistently outperforms standard Prefix-Tuning and achieves generalization capability on par with LoRA. Empirical results validate AAP as an effective, broadly applicable paradigm for parameter-efficient fine-tuning (PEFT), establishing its potential as a next-generation mainstream PEFT framework.
π Abstract
Parameter-Efficient Fine-Tuning (PEFT) methods have become crucial for rapidly adapting large language models (LLMs) to downstream tasks. Prefix-Tuning, an early and effective PEFT technique, demonstrated the ability to achieve performance comparable to full fine-tuning with significantly reduced computational and memory overhead. However, despite its earlier success, its effectiveness in training modern state-of-the-art LLMs has been very limited. In this work, we demonstrate empirically that Prefix-Tuning underperforms on LLMs because of an inherent tradeoff between input and prefix significance within the attention head. This motivates us to introduce Prefix-Tuning+, a novel architecture that generalizes the principles of Prefix-Tuning while addressing its shortcomings by shifting the prefix module out of the attention head itself. We further provide an overview of our construction process to guide future users when constructing their own context-based methods. Our experiments show that, across a diverse set of benchmarks, Prefix-Tuning+ consistently outperforms existing Prefix-Tuning methods. Notably, it achieves performance on par with the widely adopted LoRA method on several general benchmarks, highlighting the potential modern extension of Prefix-Tuning approaches. Our findings suggest that by overcoming its inherent limitations, Prefix-Tuning can remain a competitive and relevant research direction in the landscape of parameter-efficient LLM adaptation.