Quasi Linear Kernel Attention with Infinite Capacity

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the quadratic computational complexity of Transformer attention mechanisms with respect to sequence length by proposing a quasi-linear kernel attention method with infinite capacity. The authors introduce a capacity metric to quantify model expressiveness and construct additive kernel functions based on splines and polynomial exponents, thereby overcoming the limitations of finite-dimensional feature maps. Efficient computation is achieved through sorting algorithms combined with CUDA acceleration to optimize inference speed. Experimental results demonstrate that the proposed approach significantly reduces computational costs on long-sequence benchmarks while outperforming existing Softmax-based backends, effectively unifying accuracy, efficiency, and expressiveness.
πŸ“ Abstract
The evaluation cost of transformers with softmax attention scales quadratically with sequence length. Kernel attention addresses this by replacing softmax with a more general kernel function. In this paper, we aim to identify kernels that retain the expressivity of attention while enabling quasi linear computation. To quantify expressivity, we introduce a capacity for each kernel, measuring the maximum sequence length for which the attention matrix can approximate the identity. A higher capacity thus indicates greater expressivity. We show that expressive kernels like softmax, Gauss, and Laplace have infinite capacity. In contrast, common quasi linear kernels, such as those derived from finite dimensional feature maps, exhibit finite capacity. As a solution, we propose additive kernels constructed from univariate spline and polynomial exponential kernels. We prove that these maintain infinite capacity while allowing quasi linear computation via sorting. Finally, we implement additive sorting kernels efficiently and benchmark them against modern softmax backends, demonstrating advantages for long sequences.
Problem

Research questions and friction points this paper is trying to address.

kernel attention
transformers
expressivity
quasi linear computation
capacity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quasi Linear Kernel Attention
Infinite Capacity
Additive Kernels
Sorting-based Computation
Expressivity
πŸ”Ž Similar Papers
No similar papers found.