π€ AI Summary
This study addresses the challenge of disentangling the independent effects of optimizers, architectures, and data streams on performance during large language model (LLM) pretraining. To this end, it introduces the concept of Relative Generalization Invariance (RGI). Through theoretical analysis, validation using overparameterized quadratic models, and empirical investigations across diverse optimizers and architectures, the authors demonstrate that this phenomenon cannot be fully explained by either Neural Tangent Kernel (NTK) or mean-field theory alone. The findings reveal that variations in optimizers and architectures merely induce a uniform shift in loss, whereas data streams substantially alter relative generalization. By establishing RGI as a critical metric for distinguishing component-wise influences, this work significantly advances the mechanistic understanding of LLM pretraining dynamics.
π Abstract
Large Language Model (LLM) pretraining performance is jointly shaped by three components of the training triplet: the optimizer, model architecture, and training data stream. However, how these components influence performance in distinct ways remains unclear. We take a first step toward isolating their effects by studying relative generalization. We introduce Relative Generalization Invariance (RGI), the invariance of the validation-loss difference between any two tokens across models. We show that RGI approximately holds across a wide range of optimizers and moderate architectural variations, suggesting that these choices induce an approximately uniform shift in token-wise losses. In contrast, changing the training data stream can substantially alter relative generalization. We further show that RGI cannot be explained by the neural tangent kernel or mean-field regimes alone and prove that it can emerge in an overparameterized quadratic model. Overall, our work identifies RGI as a new phenomenon in LLM pretraining that helps distinguish the effects of optimizers and architectures from those of training data.