🤖 AI Summary
This study addresses the primacy bias in sequence pretraining, wherein small-scale models exhibit deficient learning capacity for later-stage data. We demonstrate that the superiority of large models stems primarily from their robustness to such training side effects rather than solely from feature representativeness. By quantitatively analyzing the detrimental effects of this bias, we propose Exposure Therapy, a regularization technique designed to optimize learning capacity allocation under heterogeneous data distributions. This approach significantly enhances the performance of sub-billion-parameter foundation models on both later-stage data and overall downstream tasks, partially recovering the performance advantages typically associated with larger architectures. Ultimately, this work establishes a novel paradigm for the efficient pretraining of small language models.
📝 Abstract
Large neural networks often acquire capabilities that small models fail to learn. Does this stem from large models learning more representative features, or from being more robust to unaccounted-for adverse effects introduced during training? We define and quantify one such adverse effect, primacy bias, as the extent to which exposure to early data distributions impairs later learning. We show that small models can allocate learning capacity inefficiently toward early distributions, whereas sufficiently overparameterized models are robust to this effect. This inefficiency is particularly consequential in pretraining, where foundation models often encounter heterogeneous data distributions sequentially rather than jointly. As a result, small foundation models can struggle to learn distributions encountered late in training, which is particularly harmful when later data emphasizes desirable capabilities such as code, mathematics, and reasoning. Motivated by these findings, we introduce Exposure Therapy (ET), a simple regularization that promotes more efficient allocation of learning capacity during sequential pretraining. We demonstrate that ET improves foundation models' performance on late data distributions as well as overall capability in models up to the billion-parameter scale. Overall, our results suggest that some benefits of large foundation models may arise from greater robustness to adverse training effects, rather than from learning more representative features, and that improved training algorithms can recover some of these advantages in smaller models.