🤖 AI Summary
This study addresses the limited long-context performance in hybrid attention models, where a seesaw effect exists between full attention and linear or sliding-window attention, and discrepancies in positional inductive biases readily induce a "short-context trap." To overcome these challenges, this work proposes a sliding-window linear attention mechanism. By integrating a hybrid attention architecture with continual pre-training and a Needle-in-a-Haystack (NIAH) evaluation framework, it systematically optimizes long-context modeling. The proposed approach achieves 16× training-free length extrapolation while maintaining 100% retrieval accuracy at a 64K context length. Ultimately, this research establishes a novel paradigm for efficient long-range reasoning in hybrid models, effectively reconciling computational efficiency with robust long-context capabilities.
📝 Abstract
The architectural design of Large Language Models (LLMs) is shifting from traditional full-attention-only models to hybrid models, which combine different attention modules to improve long-context efficiency and performance in length extrapolation and context extension. To explain why hybrid models work and how to design them better, we propose Mechanics of Long-Context Hybrid Models. As Part 1.1 of this series, we begin with hybrids of full attention and either sliding-window attention (SWA) or gated variants of linear attention (LA), represented by GLA and GDN. We first observe a Seesaw Effect in Context Extension: LA hybrids benefit more from long-context continual pretraining, whereas SWA hybrids perform better under length extrapolation. We attribute this behavior to differences in the positional inductive biases induced by these attention mechanisms. We find that SWA hybrids suffer from a Short-Context Learning Trap, Short-Window Weariness, and Long-Window Laziness, and require extended windows to enhance performance in continual long-context pretraining. For LA hybrids, we summarize the Matthew Effect of Hybrid Position Extrapolation and propose Sliding-Window Linear Attention, achieving 16$\times$ training-free length extrapolation while maintaining 100\% accuracy on NIAH-SK1 in 64k context length.