🤖 AI Summary
This study addresses the limitations of existing sparse attention methods, which treat the attention matrix as an unstructured set of values and rely on top-k selection strategies, thereby neglecting its inherent structural properties and incurring significant approximation errors. To overcome this, we propose MASA, a framework that reformulates sparse attention as a matrix approximation problem. By introducing a closed-form scoring mechanism to optimize sparse element selection for minimizing product error, MASA rectifies the theoretical shortcomings of conventional heuristic strategies. Notably, the framework operates as a plug-and-play solution without requiring modifications to underlying kernels. Extensive experiments across multiple benchmarks and large language model backbones demonstrate that MASA consistently improves inference accuracy, validating the effectiveness of the matrix approximation perspective for sparse attention.
📝 Abstract
Large Language Models (LLMs) achieve strong performance across many domains, but their efficiency is limited by the quadratic cost of attention with respect to prompt length. Sparse attention reduces this cost by retaining only a small fraction of query-key interactions to approximate the full attention matrix. However, existing methods are trapped in a mathematically wrong view: they simply keep large scalar entries or high-mass regions of the attention matrix. This treats the attention matrix as a bag of values, ignoring that it is used as a structured matrix whose entries jointly determine the attention output through multiplication with value vectors. We argue that this is the core conceptual issue: sparse attention should be formulated as matrix approximation, not as blindly choosing the largest values from a bag of entries. Based on this view, we propose Matrix Approximation Sparse Attention (MASA). MASA replaces raw attention-mass ranking with a closed-form score that measures how much each sparse unit reduces matrix-product approximation error. As a theory-grounded plug-in correction, MASA can be added to existing sparse attention frameworks without changing their sparse kernels or budgets. Extensive experiments across multiple sparse attention methods, benchmarks, and LLM backbones show consistent accuracy gains, supporting both MASA and the matrix-approximation view of sparse attention.