Score
Designs and evaluates computational approximations of the softmax and natural-exponential functions used to compute attention weights, including piecewise-linear and other exp-approximation schemes. Builds algorithms and error analyses that preserve pre-trained attention scale/temperature and task accuracy while reducing numerical or computational cost and avoiding model-specific recalibration.
The quadratic time and memory complexity of Transformer self-attention severely hinders efficient long-context modeling. This paper presents a systematic survey and reconstruction of efficient attention mechanisms for large language models, proposing the first unified taxonomy encompassing both linearization paradigms (e.g., kernel-based approximations and fast weight dynamics) and sparsification paradigms (e.g., fixed patterns, block-wise routing, and clustering-driven selection). It innovatively integrates algorithmic design with hardware-aware optimization, clarifying integration pathways for purely efficient attention and hybrid architectures in large-scale pretraining. Furthermore, it establishes a comprehensive reference framework spanning theoretical analysis, algorithmic implementation, and engineering deployment. The work delivers a systematic design paradigm and practical guidelines for scalable long-context language models.
This work addresses the performance bottleneck posed by Softmax computation in Transformer multi-head attention (MHA) modules on low-precision edge devices, particularly pronounced in small models and integer-based inference scenarios. To mitigate this, the authors propose Head-Calibrated Clipped-Linear Softmax (HCCS), a bounded monotonic alternative to the exponential function that applies clipped linear mapping to centered attention logits and incorporates lightweight, per-head calibration parameters optimized offline to preserve the original statistical properties. HCCS enables the first native int8 implementation on AMD Versal AI Engines and, when combined with quantization-aware retraining and hardware-aware co-optimization, achieves significantly higher inference throughput than existing bfloat16 or lookup-table-based approaches while maintaining task accuracy.
This study addresses the quadratic computational complexity bottleneck of the Softmax attention mechanism by investigating whether sub-quadratic time algorithms can provide non-trivial uniform approximation guarantees for all inputs under standard complexity assumptions. Employing computational complexity theory and lower bound analysis techniques combined with polynomial preprocessing, this work rigorously establishes that even with preprocessing mechanisms, achieving sub-quadratic time approximations with non-trivial uniform guarantees remains infeasible. As the first theoretical proof of this fundamental limitation, the paper delineates the computational boundaries of uniform attention approximation, providing a solid theoretical foundation for the impossibility of efficient approximation algorithms in this context.
This work challenges the conventional view that Softmax in Transformer attention is indispensable due to its probabilistic interpretation, arguing instead that its empirical success stems from implicit Frobenius-norm regularization of the attention matrix, enhancing training stability. Method: The authors theoretically establish that polynomial activation functions—without requiring non-negativity, normalization, or sparsity constraints—can equivalently enforce this norm-based regularization while preserving convergence and generalization guarantees. Their approach comprises (i) matrix-norm-theoretic modeling of attention, (ii) design of polynomial attention kernels, and (iii) end-to-end integration into standard Transformers. Results: Experiments on language modeling and machine translation show that the proposed method matches Softmax-based baselines in accuracy, improves training stability, reduces inference latency by 12%, and—critically—demonstrates, for the first time, both theoretically and empirically, the feasibility and superiority of non-probabilistic attention mechanisms.
This paper addresses a novel “rescaled softmax regression” problem arising in the attention mechanisms of large language models: minimizing the squared loss between the inner product of an exponential or hyperbolic sine/cosine function and a target vector **b**, and a nonstandard normalization term whose denominator excludes the exponential/hyperbolic function itself. We propose the first iterative algorithm applicable to multiple classes of hyperbolic functions, ensuring global convergence and computational efficiency. Theoretically, we establish that the rescaling structure guarantees robustness against input perturbations in in-context learning and derive tight stability bounds for the solution. Empirically, our modeling and optimization framework significantly enhances the robustness and generalization capability of attention modules under dynamic reasoning settings.
This work investigates the conditions under which soft-attention Transformers can exactly simulate hard attention—i.e., deterministically attending to specific subsequences of the input. Method: We introduce a synergistic mechanism combining temperature scaling with unbounded positional encodings, enabling precise control over attention concentration. Contribution/Results: We establish, for the first time, necessary and sufficient theoretical conditions for soft attention to emulate hard attention. We prove that this mechanism allows Softmax attention to exactly compute a broad class of linear temporal logic (LTL) formulas and to strictly simulate all uniform-tieless average-based hard attention models. Our analysis reveals the critical role of the temperature parameter and positional encoding design in governing logical expressivity and its transfer across attention paradigms. This significantly extends the theoretical expressive capacity of soft-attention models and provides novel foundations for interpretability and formal verification of attention mechanisms.
This study addresses the challenge that quantizing Softmax during low-precision Transformer pre-training disrupts both forward computation and backward gradient propagation, while existing calibration strategies struggle to balance efficiency and accuracy. To overcome this, we propose a quantization scheme based on K-interval attention approximation, systematically optimizing grid calibration, rounding methods, and straight-through estimator (STE) placement. Specifically, we introduce a fixed-window calibration combined with a post-normalization STE, accompanied by rigorously derived backpropagation rules. Evaluated on a 124-million-parameter model, our approach incurs only a 0.004-nat increase in validation loss at K=16, achieving near-full-precision training performance with minimal overhead. This work provides a reliable paradigm for efficient quantized pre-training.
This study addresses the integral divergence and model collapse issues arising from exponential attention in Transformers under heavy-tailed distributions. We propose replacing standard softmax with slow-growing kernels and introducing Symlog data preprocessing. By constructing a dedicated benchmark and employing Wasserstein distance metrics, we systematically evaluate the synergistic effects of slow-growing kernels and data transformations on operator learning within post-normalization architectures. Experimental results demonstrate that the proposed method effectively overcomes integral collapse, significantly outperforming conventional softmax attention mechanisms. This work establishes a theoretically grounded and practically valuable new paradigm for operator learning on heavy-tailed data.
This work investigates an exact quantum implementation of the Softmax attention mechanism under the constraint that inputs and outputs lie on a probability simplex. By leveraging amplitude encoding, Hadamard tests, and measurements via the Born rule, the computation of attention scores and value aggregation is fully mapped onto a quantum circuit, establishing for the first time an exact bijection between Softmax attention and quantum measurement. The core contributions include a unified representation of all learnable parameters as rotation gate angles, a discretized quantum interpretation of the temperature parameter, support for sparse boundary attention, and integration of techniques such as block encoding, column-loading channels, and quantum singular value transformation. An exact attention layer is realized in the infinite-sampling limit, while its fully coherent variant achieves ε-approximation with infinite circuit depth, requiring only a single measurement-and-reload step per attention score. Theoretical correctness is formally verified in Lean 4.
This study addresses the computational bottleneck of Softmax exponentiation in large language model inference by proposing Rowmax-PoT, a method that approximates exponential operations using powers of two via coarse-grained logarithmic weight representations anchored to row maxima. Motivated by the finding that resolution budget allocation matters more than absolute magnitude, we implement hardware-specialized kernels within the FlashAttention-4 framework, leveraging NVIDIA B200 architecture and FP8/BF16 mixed-precision Tensor Cores. Experiments demonstrate that on the B200 platform, forward-pass throughput for 8K sequences improves by 12.4% and 25.8% for causal and non-causal attention, respectively, while energy consumption for 16K sequences decreases by 8.4%. Notably, perplexity incurs only a marginal increase of 0.09%–0.49%, achieving substantial energy efficiency gains with minimal precision cost.
This work uncovers the fundamental reason why linearized attention mechanisms fail to converge within the Neural Tangent Kernel (NTK) framework and elucidates their dual impact on model performance and robustness. By constructing a linearized attention operator that exactly corresponds to the data-dependent Gram kernel, and integrating NTK theory with spectral analysis, the study establishes—for the first time—a theoretical link between the cubic amplification of the Gram matrix condition number and the required network width (m = Ω(κ⁶)). It introduces the notion of “influence plasticity” to unify the explanation of attention’s expressive power and its vulnerability to adversarial perturbations. Empirical results show that practical training widths fall far below the theoretical threshold, yet exhibit 6–9 times higher influence plasticity than ReLU networks, conferring superior task adaptability at the cost of heightened sensitivity to data poisoning.