length-aware attention

Design, build, or analyze attention mechanisms and softmax-normalization adjustments that account for the number or length of attention targets so that relevant items retain proportional probability mass rather than being diluted across many items. This includes methods for calibrating attention scores, modifying the softmax or its temperature, or otherwise normalizing attention to enable robust retrieval and decision-making under varying corpus sizes or extreme extrapolation.

length-awareattention

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Limitations of Normalization in Attention Mechanism

Aug 25, 2025
TM
Timur Mudarisov
🏛️ University of Luxembourg | London Institute for Mathematical Sciences

Softmax normalization in attention mechanisms exhibits fundamental limitations: it tends to distribute attention weights uniformly when selecting multiple tokens, degrading discriminative capacity; moreover, low-temperature scaling exacerbates gradient sensitivity, causing training instability. Method: We establish, for the first time, theoretical bounds on attention selection capability—grounded in vector distance metrics and geometric separation criteria—to rigorously characterize how normalization affects multi-token selection performance. We validate our analysis theoretically and empirically using GPT-2. Contribution/Results: Our analysis confirms attention degradation under high token-selection throughput and identifies this phenomenon as an intrinsic cause of training instability. The work highlights the need for more robust and scalable normalization mechanisms, providing both a novel theoretical framework and empirical evidence to guide the design of improved attention mechanisms in large language models.

Analyzes model's declining token distinction abilityExamines gradient sensitivity challenges during trainingInvestigates normalization limitations in attention mechanisms

This work proposes Affine-Scaled Attention, a novel attention mechanism that addresses the limitations of standard softmax-based attention in Transformers. The conventional softmax enforces that attention weights sum to one, which restricts flexible modulation of attention magnitudes and can lead to training instability or overly concentrated attention distributions. To overcome this, the proposed method introduces input-dependent scaling factors and bias terms after softmax normalization, enabling lightweight affine transformations that flexibly reweight attention outputs while preserving the model’s capacity for value aggregation. Experiments across Transformer models of varying scales demonstrate that Affine-Scaled Attention significantly improves training stability, optimization dynamics, and downstream task performance, consistently outperforming both standard softmax attention and recent alternatives such as Attention Sink.

attention concentrationattention flexibilityattention stability

This study addresses the challenge that quantizing Softmax during low-precision Transformer pre-training disrupts both forward computation and backward gradient propagation, while existing calibration strategies struggle to balance efficiency and accuracy. To overcome this, we propose a quantization scheme based on K-interval attention approximation, systematically optimizing grid calibration, rounding methods, and straight-through estimator (STE) placement. Specifically, we introduce a fixed-window calibration combined with a post-normalization STE, accompanied by rigorously derived backpropagation rules. Evaluated on a 124-million-parameter model, our approach incurs only a 0.004-nat increase in validation loss at K=16, achieving near-full-precision training performance with minimal overhead. This work provides a reliable paradigm for efficient quantized pre-training.

approximate softmaxgradient interactionlow-precision attention

Self-Adjust Softmax

Feb 25, 2025
CZ
Chuanyang Zheng
🏛️ The Chinese University of Hong Kong | National University of Singapore | The University of Hong Kong | Noah's Ark Lab

To address the gradient vanishing problem of Softmax under extreme attention scores in Transformers, this paper proposes Scalable Adaptive Softmax (SA-Softmax). Its core innovation is an input-dependent linear weighting mechanism, which—through theoretical analysis—is proven to improve the lower bound of gradients and enable lossless replacement of standard attention modules. The method is both lightweight and theoretically grounded: it introduces only minimal learnable scaling parameters without additional computational overhead. Extensive experiments across multilingual benchmarks, diverse datasets, and various positional encoding schemes demonstrate that SA-Softmax consistently accelerates convergence and improves final model performance on models up to 2.7B parameters. It effectively mitigates gradient vanishing while maintaining inference latency—no increase in computational cost during deployment.

Address gradient vanishing in softmaxEnhance Transformer attention mechanismsIntegrate SA-Softmax with minor adjustments

Rethinking Attention: Polynomial Alternatives to Softmax in Transformers

Oct 24, 2024
HS
Hemanth Saratchandran
🏛️ Australian Institute of Machine Learning | University of Adelaide

This work challenges the conventional view that Softmax in Transformer attention is indispensable due to its probabilistic interpretation, arguing instead that its empirical success stems from implicit Frobenius-norm regularization of the attention matrix, enhancing training stability. Method: The authors theoretically establish that polynomial activation functions—without requiring non-negativity, normalization, or sparsity constraints—can equivalently enforce this norm-based regularization while preserving convergence and generalization guarantees. Their approach comprises (i) matrix-norm-theoretic modeling of attention, (ii) design of polynomial attention kernels, and (iii) end-to-end integration into standard Transformers. Results: Experiments on language modeling and machine translation show that the proposed method matches Softmax-based baselines in accuracy, improves training stability, reduces inference latency by 12%, and—critically—demonstrates, for the first time, both theoretically and empirically, the feasibility and superiority of non-probabilistic attention mechanisms.

Exploring polynomial alternatives to softmax in transformersUnderstanding softmax's regularization effect on attentionValidating non-softmax attention mechanisms in applications

Latest Papers

What's happening recently
View more

This work investigates an exact quantum implementation of the Softmax attention mechanism under the constraint that inputs and outputs lie on a probability simplex. By leveraging amplitude encoding, Hadamard tests, and measurements via the Born rule, the computation of attention scores and value aggregation is fully mapped onto a quantum circuit, establishing for the first time an exact bijection between Softmax attention and quantum measurement. The core contributions include a unified representation of all learnable parameters as rotation gate angles, a discretized quantum interpretation of the temperature parameter, support for sparse boundary attention, and integration of techniques such as block encoding, column-loading channels, and quantum singular value transformation. An exact attention layer is realized in the infinite-sampling limit, while its fully coherent variant achieves ε-approximation with infinite circuit depth, requiring only a single measurement-and-reload step per attention score. Theoretical correctness is formally verified in Lean 4.

Born ruleprobability simplexquantum attention

This work addresses the lack of a unified theoretical understanding of how the inverse temperature should scale with context length in long-context self-attention mechanisms. The authors propose a unified framework based on an upper-tail cumulative scale derived from a gap-counting function of attention score rows, which determines the critical inverse temperature scaling required for softmax concentration. Through probabilistic concentration analysis and investigation of attention entropy dynamics, they demonstrate that this scale marks the threshold between entropy collapse and well-behaved attention distributions. The resulting theory, which depends on the number of competing gaps, unifies several previously proposed scaling laws. The framework applies broadly—from idealized models to practical Transformers—and yields a readily applicable diagnostic tool for practitioners.

context lengthinverse temperaturescaling law

This study investigates whether the "attention collapse" phenomenon in Softmax-based self-attention is inevitable and reveals its intrinsic connection to normalization constraints. By designing a trigger-condition task—where the model outputs the mean of historical representations only upon encountering a specific trigger token and zero otherwise—the authors theoretically demonstrate that Softmax normalization necessarily induces attention collapse to implement the default (zero) behavior. In contrast, non-normalized ReLU attention entirely avoids this issue. Through comparative experiments on synthetic tasks and extended scenarios using both single-head and multi-head architectures, the work empirically validates that normalization is the root cause of attention collapse, provides the first theoretical justification for its necessity under Softmax, and demonstrates the effectiveness of ReLU attention in eliminating such collapse.

attention sinknormalizationself-attention

This work addresses the pervasive position bias in dense retrieval models, which significantly degrades recall performance for relevant content located toward the end of passages. The authors propose a training-free, inference-time attention calibration mechanism that interpolates attention weights using an adjustable strength coefficient λ. This approach integrates hierarchical calibration and basket sampling strategies, making it compatible with both <s>-token pooling and last-token pooling architectures. Evaluated under a unified default configuration, the method consistently enhances positional fairness across diverse models, architectures, and languages while preserving or even improving overall retrieval effectiveness. Specifically, it substantially increases the harmonic mean of nDCG@10 across position groups on FineWeb-PosQ and comprehensively reduces the position sensitivity index on the multilingual, multidomain PosIR benchmark.

attention calibrationdense retrievalinference-time adaptation

This study investigates the degenerate behavior of softmax self-attention in long-context Transformers as the context length tends to infinity. Focusing on the setting where queries are fixed and keys are random, the authors introduce an inverse temperature parameter to systematically characterize a phase transition in attention—from uniform averaging to collapse onto a single key. By integrating tools from probability limit theory, spherical geometry, and random matrix theory, they show that the critical scale governing attention selectivity is determined by the local exponent of the query-key distance distribution. Notably, in the subcritical regime, the attention map approximates the backward heat equation. The work fully establishes the limiting distributions of attention outputs across subcritical, critical, and supercritical regimes, capturing phenomena such as Gaussian fluctuations, finite-neighbor mass retention, and point concentration.

attention degeneracyinverse temperaturelong-context transformers

Hot Scholars

CW

Chunwei Wang

Researcher, Huawei Noah's Ark Lab
Computer VisionAutonomous DrivingMultimodality
WK

Wonbin Kweon

University of Illinois Urbana-Champaign
Data MiningText MiningRecommender SystemsInformation Retrieval
WZ

Wangmeng Zuo

School of Computer Science and Technology, Harbin Institute of Technology
Computer VisionImage ProcessingGenerative AIDeep Learning
YY

Yitao Yang

University of Leeds
Human MobilityUrban AnalyticsComputational Social ScienceComplex Systems