Dynamic Positional Attention Modulation for Parameter-Efficient Fine-Tuning of Large Language Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing parameter-efficient fine-tuning (PEFT) methods in adapting to the positional encoding characteristics induced by Rotary Position Embedding (RoPE), primarily due to their neglect of attention structural heterogeneity. To this end, this work proposes DyPAM, a method that dynamically modulates positional information within query and key representations without modifying the pretrained backbone. Specifically, DyPAM achieves fine-grained, structured attention adaptation through a synergistic mechanism combining input-conditioned dimensional modulation with head- and layer-wise structural modulation. By directly aligning with the RoPE positional structure, the proposed approach significantly outperforms strong existing baselines across multiple models on mathematical and commonsense reasoning benchmarks.
📝 Abstract
Parameter-efficient fine-tuning (PEFT) has become a standard approach for adapting large language models to downstream tasks. However, most existing PEFT methods rely on uniform and static adaptations, without accounting for the structured heterogeneity of attention across dimensions, heads, layers, and input tokens. In practice, attention representations exhibit non-uniform behavior, and positional encoding mechanisms such as rotary positional embeddings (RoPE) induce dimension-dependent positional structure, making uniform adaptation suboptimal. In this work, we propose DyPAM (Dynamic Positional Attention Modulation), a PEFT method that adapts how positional information contributes to attention by operating directly on the query and key representations. DyPAM combines input-conditioned, dimension-wise modulation with head-wise and layer-wise structural modulation, performing fine-grained adaptation of positional attention aligned with the RoPE-induced structure without modifying the pretrained backbone. Extensive experiments on mathematical and commonsense reasoning benchmarks across multiple backbone models demonstrate that DyPAM consistently outperforms existing strong PEFT baselines.
Problem

Research questions and friction points this paper is trying to address.

Parameter-Efficient Fine-Tuning
Large Language Models
Positional Attention
Rotary Positional Embeddings
Structured Heterogeneity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parameter-Efficient Fine-Tuning
Dynamic Positional Attention Modulation
Rotary Positional Embeddings
Dimension-wise Modulation
Large Language Models
D
Dayan Pan
School of Computer Science and Engineering, MOE Engineering Research Center of Advanced Computer Application Technology, Beihang University, Beijing, China
Jingyuan Wang
Jingyuan Wang
Beihang University
Data MiningSpatio-temporal Data MiningUrban ComputingCOVID-19
X
Xie Yu
School of Computer Science and Engineering, MOE Engineering Research Center of Advanced Computer Application Technology, Beihang University, Beijing, China