Emergent One-Third Scaling Law as Attention Tries to Concentrate

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear origin of the 1/3 neural scaling law governing loss scaling in large language models (LLMs). By integrating theoretical modeling, toy experiments, and large-scale empirical validation, this work provides the first theoretical proof that when any softmax function learns a spiky distribution, its logit magnitudes grow according to a 1/3 power law, thereby driving the overall loss to follow this scaling behavior. Furthermore, the analysis reveals that the attention mechanism constitutes the critical bottleneck underlying this phenomenon. These findings demonstrate that the primary limitation on LLM training performance stems from attention concentration rather than the language modeling head. Ultimately, this research establishes a novel theoretical foundation for understanding and overcoming the scaling limits of large language models.
📝 Abstract
The neural scaling law relating longer training to better performance through a power law is central to today's large language models (LLMs), yet its origin remains debated. One recent proposal is that power laws can emerge from the strong non-linearity of a single softmax head learning peaked distributions. What happens with multiple softmax functions, as in LLMs, is unclear. Here, we show through toy models that any softmax learning peaked distributions, regardless of its position in the model, can develop logit magnitudes that grow in a power law with exponent $1/3$, becoming a training bottleneck whose loss contribution decays as a power law with the same exponent $1/3$. The overall loss therefore obeys $1/3$ scaling whenever at least one softmax learns peaked distributions. We confirm that many softmax functions in LLMs learn peaked distributions and that LLM loss scaling matches this $1/3$ prediction. Moreover, logit growth dynamics reveal that attention heads, rather than the language modeling head, are the bottleneck likely driving the $1/3$ loss scaling in LLMs. Attention trying to concentrate on specific information, which is the heart of Transformers, may therefore also be the heart of the neural scaling law of training.
Problem

Research questions and friction points this paper is trying to address.

Neural Scaling Law
Large Language Models
Softmax
Attention Mechanism
Power Law
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scaling Law
Softmax
Attention Mechanism
Power Law
Large Language Models
🔎 Similar Papers
No similar papers found.