Efficiently Approximating Attention Is Hard

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the quadratic computational complexity bottleneck of the Softmax attention mechanism by investigating whether sub-quadratic time algorithms can provide non-trivial uniform approximation guarantees for all inputs under standard complexity assumptions. Employing computational complexity theory and lower bound analysis techniques combined with polynomial preprocessing, this work rigorously establishes that even with preprocessing mechanisms, achieving sub-quadratic time approximations with non-trivial uniform guarantees remains infeasible. As the first theoretical proof of this fundamental limitation, the paper delineates the computational boundaries of uniform attention approximation, providing a solid theoretical foundation for the impossibility of efficient approximation algorithms in this context.
📝 Abstract
Softmax attention is ubiquitous in modern machine learning, but its quadratic scaling with sequence length makes it costly. To reduce this cost, attention is often approximated with fast algorithms, which incur error but can still perform well in practice and on some inputs. At the same time, the growing diversity of attention applications makes approximation guarantees that do not depend on particular input structure a compelling target. For such uniform guarantees over all inputs, known runtime lower bounds rule out fast algorithms for near-exact attention, but leave open the practically important regime: is there an efficient algorithm with even a modest uniform approximation guarantee? We answer this question negatively. Under standard complexity-theoretic assumptions, no truly subquadratic algorithm can approximate attention with any nontrivial additive or relative guarantee uniformly over all inputs. This impossibility holds in the mildest parameter regime for which known algorithms do not already achieve strong approximation guarantees in near-linear time, and extends to practically relevant relaxations: even after polynomial preprocessing of the KV cache, no efficient algorithm can obtain a nontrivial uniform approximation guarantee, or identify a small set of keys receiving substantial attention under sparsity. Overall, our results settle the computational limits of uniform attention approximation.
Problem

Research questions and friction points this paper is trying to address.

Attention Approximation
Computational Complexity
Subquadratic Algorithms
Softmax Attention
Approximation Guarantees
Innovation

Methods, ideas, or system contributions that make the work stand out.

Attention Approximation
Computational Complexity
Subquadratic Algorithms
KV Cache
Lower Bounds
🔎 Similar Papers
No similar papers found.