How Local Mixing Encodes Relative Position in Global NoPE Attention

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how global attention implicitly acquires relative positional information in hybrid models without explicit position encoding (NoPE). By integrating theoretical derivations with empirical analyses, we elucidate the mechanism through which recency biases induced by local mixture layers within the residual stream propagate to global attention. To our knowledge, this work provides the first mechanistic explanation of positional encoding principles in NoPE hybrid architectures, demonstrating their capacity to effectively preserve positional information over long sequences while outperforming purely global NoPE counterparts. These findings substantially advance the understanding of positional representations in hybrid models and establish a critical theoretical foundation for developing novel position encoding methods capable of supporting infinite-length extrapolation.
📝 Abstract
The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of explicit position encodings have been developed in transformer-based models, such as rotary position encoding (RoPE). Although explicit position encodings have long been assumed to be required, recent methods that interleave local mixing layers, such as sliding window attention (SWA) and gated linear attention, while not encoding position (NoPE) in global attention layers has recently been shown to be successful at scale. How and why this approach works is not well-understood. In this paper, we develop an explanation of how hybrid models of this sort can implicitly encode position at global NoPE layers. Supported by both theoretical and empirical evidence, our central argument is that SWA and gated linear attention induce a recency bias in the residual stream that propagates to, and is selected by, the global attention logits. Moreover, in contrast to the implicit position encodings found in models with only global NoPE attention, in which positional information arises solely from the causal mask, the recency bias in hybrid models can be maintained across long sequences. In addition to deepening our understanding of how hybrid models encode position, these findings may provide insights for how to encode position in a way that can extrapolate to longer sequence lengths indefinitely.
Problem

Research questions and friction points this paper is trying to address.

Position Encoding
NoPE Attention
Hybrid Models
Sliding Window Attention
Recency Bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

NoPE Attention
Sliding Window Attention
Recency Bias
Implicit Position Encoding
Hybrid Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Cutter Dawes
Zyphra Research
N
Nick Alonso
Zyphra Research
T
Tom Figliolia
Zyphra Research
Beren Millidge
Beren Millidge
Postdoctoral Researcher, University of Oxford