🤖 AI Summary
This study addresses the bottleneck in hybrid architectures where exact attention is confined to fixed local windows, hindering the simultaneous modeling of long-range dependencies and historical context. To this end, we propose a Prefix-State Mixture Block Attention mechanism. This method replaces sliding windows with top-k block-sparse retrieval and innovatively integrates gated linear recurrence to construct compressed prefix states, thereby unifying long-range exact retrieval with global historical context modeling. Furthermore, Triton hardware-aware operators are employed to stream-process routed blocks and prefix states, avoiding the materialization of large intermediate tensors. Experiments demonstrate that the proposed mechanism significantly enhances long-context understanding and information retrieval performance while maintaining efficient training and inference, outperforming existing baseline models.
📝 Abstract
Hybrid architectures combining linear sequence models with softmax attention provide an effective balance between efficient long-context modeling and precise token retrieval. Existing designs such as Native Hybrid Attention (NHA) combine compressed long-term states with sliding-window attention, but their exact attention is restricted to a fixed local window. In this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context. The prefix states are constructed by a gated linear recurrence at block boundaries and retrieved together with the corresponding token blocks, allowing the model to combine precise long-range evidence with compressed historical context within a unified layer. We further develop a hardware-aware Triton implementation that streams routed token blocks and prefix states without materializing large intermediate tensors. Experiments show that PHBA improves long-context and retrieval performance over strong linear and hybrid baselines while retaining efficient training and inference.