PHBA: Prefix-State Hybrid Block Attention

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottleneck in hybrid architectures where exact attention is confined to fixed local windows, hindering the simultaneous modeling of long-range dependencies and historical context. To this end, we propose a Prefix-State Mixture Block Attention mechanism. This method replaces sliding windows with top-k block-sparse retrieval and innovatively integrates gated linear recurrence to construct compressed prefix states, thereby unifying long-range exact retrieval with global historical context modeling. Furthermore, Triton hardware-aware operators are employed to stream-process routed blocks and prefix states, avoiding the materialization of large intermediate tensors. Experiments demonstrate that the proposed mechanism significantly enhances long-context understanding and information retrieval performance while maintaining efficient training and inference, outperforming existing baseline models.
📝 Abstract
Hybrid architectures combining linear sequence models with softmax attention provide an effective balance between efficient long-context modeling and precise token retrieval. Existing designs such as Native Hybrid Attention (NHA) combine compressed long-term states with sliding-window attention, but their exact attention is restricted to a fixed local window. In this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context. The prefix states are constructed by a gated linear recurrence at block boundaries and retrieved together with the corresponding token blocks, allowing the model to combine precise long-range evidence with compressed historical context within a unified layer. We further develop a hardware-aware Triton implementation that streams routed token blocks and prefix states without materializing large intermediate tensors. Experiments show that PHBA improves long-context and retrieval performance over strong linear and hybrid baselines while retaining efficient training and inference.
Problem

Research questions and friction points this paper is trying to address.

hybrid attention
long-context modeling
sliding-window attention
token retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid Attention
Block-Sparse Retrieval
Prefix State
Gated Linear Recurrence
Hardware-Aware Triton