🤖 AI Summary
This study addresses the lack of rigorous theoretical analysis regarding the spectral properties and structural dependencies of multi-head self-attention mechanisms. Leveraging random matrix theory and Gaussian process modeling, this work establishes the first rigorous Gaussian equivalence framework for multi-head attention, proving that the softmax operation can be replaced by a noisy rescaled score without altering the limiting spectral distribution of the output. This approach effectively decouples the effects of head allocation from projection width while distinguishing cross-head sharing from key-value dependencies. By revealing the spectral distribution patterns of attention outputs and elucidating how distinct architectural components independently influence model behavior, this research provides a solid theoretical foundation for understanding multi-head attention mechanisms.
📝 Abstract
A theoretical understanding of multi-head self-attention is fundamental to the study of modern neural networks. Using random matrix theory, we establish Gaussian equivalence for multi-head self-attention: replacing softmax attention with rescaled scores plus Gaussian noise preserves the limiting spectral law of the centered output. This equivalence also covers value and output projections that depend on the keys. The resulting laws separate the effects of head allocation and projection widths, and distinguish spectrum-preserving across-head sharing from within-head key--value dependence.