🤖 AI Summary
To address the challenge of balancing parameter efficiency and cross-task generalization in single-backbone multi-task fine-tuning, this paper proposes PARA, a lightweight parameter-efficient fine-tuning (PEFT) method. Its core innovation lies in introducing a prompt-aware vector generator into each Transformer layer—enabling fine-grained, low-overhead dynamic adaptation of hidden-layer representations. PARA further integrates intra-layer feature reweighting with strategic parameter freezing, achieving substantial gains in multi-task generalization while adding only 0.1% extra parameters—comparable to LoRA. Empirical evaluation on comprehensive multi-task benchmarks demonstrates that PARA consistently outperforms existing PEFT methods. Moreover, under single-backbone multi-tenant deployment, PARA accelerates inference by 17% relative to LoRA, confirming its superior efficiency–effectiveness trade-off.
📝 Abstract
In the realm of parameter-efficient fine-tuning (PEFT) methods, while options like LoRA are available, there is a persistent demand in the industry for a PEFT approach that excels in both efficiency and performance within the context of single-backbone multi-tenant applications. This paper introduces a new and straightforward PEFT technique, termed Prompt Aware Representation Adjustment (PARA). The core of our proposal is to integrate a lightweight vector generator within each Transformer layer. This generator produces vectors that are responsive to input prompts, thereby adjusting the hidden representations accordingly. Our extensive experimentation across diverse tasks has yielded promising results. Firstly, the PARA method has been shown to surpass current PEFT benchmarks in terms of performance, despite having a similar number of adjustable parameters. Secondly, it has proven to be more efficient than LoRA in the single-backbone multi-tenant scenario, highlighting its significant potential for industrial adoption.