Superlinear Multi-Step Attention

📅 2026-01-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high computational complexity—typically O(L²)—of standard causal self-attention over long sequences and its inability to support random access to arbitrary context positions. The authors propose a multi-step attention architecture that reformulates causal self-attention as a learnable, stepwise search process. By integrating span selection and local attention mechanisms inspired by skip-list structures, the method achieves sub-quadratic complexity of O(L^{1+1/N}) for the first time while preserving non-exclusive attention patterns and enabling random access. Implemented as a two-step attention pipeline within a 30B sparse Mixture-of-Experts (MoE) model, the approach demonstrates strong empirical performance: on a single B200 GPU, it achieves decoding speeds of 114 tokens/s at 1M context length and 80 tokens/s at 10M. Furthermore, it successfully handles 256K-context tasks in Needle-in-a-Haystack (NIAH) evaluations, confirming the learnability of its routing mechanism and effectiveness in ultra-long-context modeling.

Technology Category

Machine Learning: Mixture of Experts (MoE)Search and Optimization: Learning to SearchNatural Language Processing: (Large) Language Models

Application Category

Search and Retrieval-Augmented AI: Personalized, context-aware and across-device searchGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
In this paper, we propose \textbf{Superlinear attention}, a fully trainable multi-step attention architecture that achieves subquadratic complexity for long sequences while preserving \textbf{random context access} (a.k.a.\ structural non-exclusion): no eligible token position is structurally excluded from being selected for attention. Superlinear attention reformulates standard causal self-attention as a multi-step search problem with $N$ steps, yielding an overall complexity of $O(L^{1+\frac{1}{N}})$. To illustrate the architecture, we present a baseline $N=2$ implementation, which is algorithmically analogous to standard jump search. In this $O(L^{3/2})$ instantiation, the first step performs $O(L^{3/2})$ span-search to select relevant spans of the sequence, and the second step applies $O(L^{3/2})$ span-attention (standard attention restricted to the selected spans). In an upscaled $O(L^{1.54})$ configuration for robustness, we achieve an average decoding throughput of 114 tokens/sec at 1M context length and 80 tokens/sec at 10M context in our implementation on a modified 30B hybrid MoE model on a single B200 GPU. With limited training, we also obtain strong performance on the NIAH (Needle In A Haystack) task up to 256K context length, demonstrating that the routed span selection is learnable end-to-end. This paper emphasizes architectural formulation, scaling analysis, and systems feasibility, and presents initial validation; comprehensive quality evaluations across diverse long-context tasks are left to future work.
Problem

Research questions and friction points this paper is trying to address.

superlinear attention
subquadratic complexity
random context access
long-context attention
multi-step attention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Superlinear attention
subquadratic complexity
random context access
multi-step attention
long-context modeling
🔎 Similar Papers
No similar papers found.