Block-Sparse Attention with Semantic-Geometric Decoupled Routing

📅 2026-09-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决长文本推理中密集注意力成本高的问题,提出了一种语义-几何解耦路由方法,实现高效块稀疏注意力,减少计算开销。
📝 Abstract
Long-context inference has become a defining capability of large language models, but exact dense attention remains costly due to its quadratic scaling with sequence length. Block-sparse attention offers a hardware-friendly alternative by routing each query block to a small set of relevant key blocks, yet accurate training-free block routing remains difficult. Existing routers often pool post-RoPE token representations, which entangles semantic aggregation with RoPE-induced geometry and attenuates local positional cues through high-frequency phase cancellation. To resolve this mismatch, we propose \textbf{Semantic-Geometric Decoupled Routing}, a training-free block routing framework that shifts semantic aggregation to the pre-RoPE space and reconstructs geometric bias with an offline structural prior and relative block distances. This decomposition yields an explicit closed-form block routing score without token-level search or post-hoc calibration. Experiments on long-context text and video tasks show that our method approaches full-attention accuracy across 4K--128K contexts, keeps routing overhead below 3.4 ms, and achieves a 5.03$\times$ speedup over FlashAttn at a 128K context length.
Problem

Research questions and friction points this paper is trying to address.

Block-Sparse Attention
Semantic-Geometric Decoupled Routing
Long-context Inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Block-Sparse Attention
Semantic-Geometric Decoupled Routing
Training-Free Block Routing
🔎 Similar Papers
No similar papers found.