The Life Cycle of a Massive Activation: Stochastic Birth, Weight-Decay-Driven Growth, and Competitive Consolidation

๐Ÿ“… 2026-09-30
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study investigates the lifecycle of massive activations in Transformers and their dynamic regulation mechanisms during training. Through training trajectory analysis, controlled intervention experiments, and ablation studies, we systematically trace the emergence, growth, and consolidation of these extreme activation values. Our findings reveal a causal role of weight decay in governing global activation scales, leading us to propose a balancing model between AdamW preconditioning and decay antagonism that establishes it as a critical lever for modulating activation magnitudes. Experiments confirm that peak amplitudes scale according to a ฮปโปยน/ยฒ law with respect to the decay coefficient, elucidating the central role of collective scaling in gradient attenuation. This work provides new perspectives for understanding and controlling activation dynamics in large-scale models.
๐Ÿ“ Abstract
Massive activations, residual-stream coordinates with magnitudes far larger than typical activations, are associated with attention sinks in transformers, but how their scale is regulated during training remains incompletely understood. Combining training-trajectory analyses and controlled interventions, we trace their emergence, growth, and consolidation. Sink-carrying channels vary across random seeds but stabilize early within each run. Over longer training, surrounding channels erode and the sink concentrates onto a few redundant carriers. Across ablations, gradient attenuation follows the sink token's collective root-mean-square magnitude rather than any single channel, making collective scale central to understanding their effects. Our central result is that weight decay causally controls the turnover of global activation scale. In controlled continuations, removing decay near the peak allows this scale to keep rising, whereas retaining it produces decline even at constant learning rate. We develop a balance model for the rise and peak of massive-activation magnitude, in which AdamW-preconditioned growth opposes weight decay. Sweeping the decay coefficient $ฮป$ shifts peak timing approximately log-linearly and yields peak magnitudes scaling approximately as $ฮป^{-1/2}$, consistent with this balance. Optimizer measurements further show that preconditioning sustains the large-channel cohort against decay even when raw maintaining forces are too small to do so. Together, these findings connect the observed life cycle to scale-regulating training dynamics and establish weight decay as a training-time lever on activation magnitude.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Massive Activations
Weight Decay
Attention Sinks
Training Dynamics
AdamW Preconditioning
๐Ÿ”Ž Similar Papers
๐Ÿ’ผ Related Jobs
No related jobs found.
S
S. Aaron McClendon
Aimpoint Digital Labs, Atlanta, GA, USA
J
Jorge Gallego-Feliciano
Antonios Saravanos
Antonios Saravanos
Clinical Professor of Information Systems, New York University
Human-Computer Interaction