Score
Designs, implements, and evaluates attention modules that enforce sparse connectivity patterns—including dynamic/input-dependent masks, hierarchical or landmark-based sparsity, and chunk-wise/blockwise attention—to reduce computation and memory while preserving model output quality. Work covers profiling attention heads offline, generating and updating sparse masks online, building efficient sparse attention implementations, and analyzing trade-offs among sparsity structure, efficiency, and performance.
The quadratic time and memory complexity of Transformer self-attention severely hinders efficient long-context modeling. This paper presents a systematic survey and reconstruction of efficient attention mechanisms for large language models, proposing the first unified taxonomy encompassing both linearization paradigms (e.g., kernel-based approximations and fast weight dynamics) and sparsification paradigms (e.g., fixed patterns, block-wise routing, and clustering-driven selection). It innovatively integrates algorithmic design with hardware-aware optimization, clarifying integration pathways for purely efficient attention and hybrid architectures in large-scale pretraining. Furthermore, it establishes a comprehensive reference framework spanning theoretical analysis, algorithmic implementation, and engineering deployment. The work delivers a systematic design paradigm and practical guidelines for scalable long-context language models.
Existing sparse attention methods for mitigating the quadratic complexity bottleneck in the prefill phase of long-context large language models rely on predefined patterns or coarse approximations, failing to accurately capture real attention dynamics and thus compromising the accuracy–efficiency trade-off. This work first identifies and empirically validates strong intra-layer similarity across attention heads—a consistent structural property of multi-head attention. Leveraging this insight, we propose a cross-head sparse pattern sharing mechanism: a lightweight dynamic module extracts a shared sparse attention pattern, while preserving full attention computation for a small set of critical heads. Our method maintains state-of-the-art accuracy while significantly accelerating prefill. On multiple long-text benchmarks, it achieves prefill throughput competitive with or superior to current best sparse approaches. This establishes a new paradigm for efficient long-context inference.
This work addresses the poor interpretability and structural redundancy inherent in Transformer attention mechanisms. We propose a post-training sparsification method that optimizes pretrained attention weights under constrained sparsity regularization, incorporating structural priors to uncover intrinsic causal circuits—without degrading model performance. Our approach globally simplifies the computational circuit: attention connections are reduced to approximately 0.3% of their original count, and task-relevant computational edges decrease by up to 100×, while strictly preserving the original pretraining loss. Empirical evaluation confirms effectiveness on billion-parameter models. The core contribution lies in redefining sparsity as an interpretability-driven structural inductive bias—not merely an efficiency heuristic—enabling lossless, global circuit simplification and explicit mechanistic revelation.
To address the quadratic computational complexity of self-attention in large language models (LLMs) for long-context modeling, this paper proposes a dynamic sparse attention mechanism. Our method introduces the first learnable sparse mask jointly conditioned on content and position: it dynamically generates sparsity patterns from key-value content while incorporating a sliding window to enforce position-aware sparse computation, fully compatible with multi-head and multi-query attention architectures. Unlike static sparsity or fixed-window approaches, our mechanism adaptively concentrates computation on salient tokens and eliminates redundant operations, thus balancing efficiency and modeling fidelity. Under the Chinchilla scaling law, our 1.7B-parameter model achieves significant improvements in both accuracy and throughput over multi-head attention, sliding-window attention, and state-of-the-art sparse attention baselines—across standard long-context benchmarks and the “needle-in-a-haystack” long-range reasoning task.
Dense attention in large language models incurs O(N²H) computational complexity, severely limiting training efficiency for long contexts; existing sparse attention methods struggle to balance efficiency and modeling capacity. This paper introduces SPAttention, the first framework for *principled structural sparsity*: it partitions multi-head attention by token distance across heads, enabling functional specialization and collaborative interaction among heads—replacing independent head computations with a structured inductive bias. This design ensures balanced computational load, eliminates redundancy, and unifies multi-head attention into a coherent collaborative process. Evaluated on the OLMoE model family, SPAttention achieves nearly 2× higher training throughput while matching or surpassing dense attention in performance, and consistently outperforms established baselines—including Longformer, Reformer, and BigBird—across multiple benchmarks.
The intrinsic nature of token-wise feature interaction in Transformer attention remains poorly understood. Method: We propose Low-Rank Sparse Attention (Lorsa), the first structured and invertible decomposition of multi-head self-attention (MHSA) into interpretable low-rank (capturing global coordination) and sparse (encoding local specificity) components, enabled by dictionary learning and rigorous interpretability analysis. Contribution/Results: Applied to Llama-3.1-8B, Lorsa automatically identifies dedicated attention head families performing atomic arithmetic operations and refines canonical circuit patterns—including induction heads and successor heads. Compared to sparse autoencoders (SAEs), Lorsa achieves significantly improved circuit discovery capability while matching SAE-level interpretability. This establishes a new analytical paradigm for attention mechanisms that is both mathematically rigorous and cognitively transparent.
This work proposes a parameter-free sparse attention mechanism that circumvents the high deployment costs of existing adaptive approaches, which often rely on learnable parameters, custom gradient estimators, or specialized CUDA kernels. The method introduces a novel, parameter-free signal for content selection by leveraging gzip compression ratios: it dynamically identifies information-rich segments of text through their compressibility and constructs sparse attention masks accordingly. This approach integrates seamlessly into standard Transformer architectures without architectural modifications. Evaluated on the PG-19 dataset with an 8K-token context window, the model achieves 1.71 bits per byte (BPB), substantially outperforming both dense and various sparse baselines. It also converges 3.3 times faster, with performance gains increasing as sequence length grows.
This work addresses the lack of a unified theoretical foundation in existing attention mask designs. It establishes, for the first time, a formal connection between attention masks and partially ordered structures, proving that information flow in sufficiently deep multi-layer Transformers converges to a Hasse diagram. The mask design problem is thereby reformulated as finding the minimal common supergraph of such Hasse diagrams, yielding a general framework that derives attention masks directly from task families. Leveraging this framework, the authors propose two novel mechanisms—Block Two-Stream Attention and Butterfly Attention—and derive block-wise causal masks and fully supervised bidirectional masks that guarantee consistency between training and inference. Empirical results validate both the effectiveness and generality of the proposed approach.
This work addresses the performance bottleneck caused by KV cache loading in long-context large language model inference and the accuracy degradation stemming from existing block-sparse attention methods that employ a uniform block size across all attention heads, ignoring their varying sensitivity to block granularity. To overcome these limitations, the paper proposes a training-free algorithm-system co-design framework that introduces, for the first time, an adaptive block size allocation mechanism across attention heads. This approach integrates lossless block centroid quantization with customized GPU kernels to significantly improve inference accuracy while preserving throughput. Experimental results demonstrate that, at comparable throughput levels, the proposed method achieves up to a 5.43% higher inference accuracy compared to current block-sparse baselines.
Sparse-Linear Attention (SLA) combines sparse and linear attention to accelerate diffusion models and has shown strong performance in video generation. However, (i) SLA relies on a heuristic split that assigns computations to the sparse or linear branch based on attention-weight magnitude, which can be suboptimal. Additionally, (ii) after formally analyzing the attention error in SLA, we identify a mismatch between SLA and a direct decomposition into sparse and linear attention. We propose SLA2, which introduces (I) a learnable router that dynamically selects whether each attention computation should use sparse or linear attention, (II) a more faithful and direct sparse-linear attention formulation that uses a learnable ratio to combine the sparse and linear attention branches, and (III) a sparse + low-bit attention design, where low-bit attention is introduced via quantization-aware fine-tuning to reduce quantization error. Experiments show that on video diffusion models, SLA2 can achieve 97% attention sparsity and deliver an 18.6x attention speedup while preserving generation quality.