Shifting Mechanisms: How Positional Encoding Choice Shapes In-Context Retrieval

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how the choice of positional encoding (PE) reshapes the contextual retrieval mechanisms of large language models. Through mechanistic interpretability analysis across 22 open-source models and controlled pretraining ablations, it provides the first quantitative assessment of PE’s remodeling effects on internal retrieval. The findings reveal that hybrid encoding architectures shift retrieval from position-dependent to semantically driven processes; however, long-context performance gains obscure a degradation in competitive key discrimination. Furthermore, retaining PE in local layers is shown to induce semantic drift, while the trade-off characteristics of SWA NoPE are clarified alongside significant discrepancies among hybrid architectures in multi-objective retrieval tasks.
📝 Abstract
Language models increasingly use architectures that vary attention span and positional encoding across layers, such as applying RoPE with sliding-window attention and NoPE with global attention (SWA NoPE). However, how these choices shape in-context retrieval remains unclear. To study this question, we take a mechanistic view, tracing how positional encoding (PE) choice shapes the internal mechanisms models use for in-context retrieval. Across 22 open-weight models spanning eight families, we find that standard RoPE models rely primarily on positional retrieval, while PE hybrids shift toward semantic retrieval. We further show on a controlled pre-training ablation that confining positional encoding to local layers produces this semantic shift, degrading representations of positional information. Finally, we show that the reported long-context gains of PE hybrids mask a retrieval trade-off: SWA NoPE improves over RoPE on multiple-target retrieval and QA, but degrades when distinguishing competing keys. We show that these behavioral differences better track the mechanism shift from positional toward semantic mechanisms than a uniform improvement in long-context retrieval.
Problem

Research questions and friction points this paper is trying to address.

positional encoding
in-context retrieval
language models
mechanistic interpretability
attention mechanisms
Innovation

Methods, ideas, or system contributions that make the work stand out.

Positional Encoding
In-Context Retrieval
Mechanistic Interpretability
Rotary Position Embedding (RoPE)
Sliding Window Attention