🤖 AI Summary
This work investigates whether the attention mechanism in Transformers acts as a relevant, marginal, or irrelevant operator in the sense of the renormalization group (RG), depending on the underlying data correlation structure. By interpreting attention as a perturbation to the fixed point of an MLP residual stack, the study introduces Wilson’s RG framework into Transformer analysis for the first time, demonstrating that operator relevance is governed by the spectral structure of the data rather than architectural choices. Through RG flow analysis, effective rank profiles, and layer-specific perturbation decay spectra—validated on controlled Markov chain sequences—the authors show that attention behaves as a strongly relevant operator under long-range correlations, inducing a phase transition in representation space, but becomes irrelevant under short-range correlations. The study further reveals that the first head of the first layer dominates representational evolution and that attention exhibits opposing pattern-selection behaviors across correlation regimes, yielding four testable predictions.
📝 Abstract
Using the language of Wilsonian renormalization group theory (RG), we treat the Transformer's attention mechanism as a perturbation of the trained MLP residual-stack fixed point and ask whether it constitutes a relevant, marginal, or irrelevant operator. We derive a fixed-point shift formula and obtain four testable predictions for the fixed-point geometry, effective rank profile, layer specificity, and perturbation decay spectrum. Testing these on synthetic Markov chain sequences with controlled correlation length, we find: (1) For large chains(long correlation), attention is strongly relevant: it closes a residual loss gap the MLP cannot bridge and drives a phase transition in representation space, with effective rank jumping above input dimensionality at layer 1 and stabilizing at a high-dimensional plateau. (2) For short chains(short correlation), attention is irrelevant: the Transformer converges to the same loss and fixed-point geometry as the MLP, though it contracts perturbations faster. (3) The transition is dominated by the first-layer head (L0H0), which accounts for more than 4 times the representational shift of any subsequent head, consistent with the prediction that the relevant operator acts before the MLP begins integrating out positional variation. (4) Perturbation decay experiments reveal a regime reversal: in the long correlation regime the Transformer selectively preserves slow Markov modes (5.4 times the dynamic range in decay length vs. 1.3 times for the MLP); in the short correlation regime it suppresses all modes faster than the MLP, with no spectral selectivity. Together, these results show that the relevance of attention is not a property of the architecture but of the spectral structure of the data-generating process, and that a first-order RG perturbation framework provides a predictive account of that difference.