What Is a Repeated Token Worth? The Scaling Geometry of Multi-Epoch Pretraining

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses how to jointly optimize model size, training epochs, and data repetition strategies to enhance pre-training efficiency under data-constrained scenarios. It proposes a "repetition cost" theory that establishes a scaling law based on unique tokens per parameter by contrasting the value of fresh versus repeated tokens, systematically validated through large-scale language model experiments, loss analysis, and entropy evaluation. The findings reveal that the critical number of training epochs is determined by the compute budget rather than model scale. Furthermore, this work quantifies the impact of different repetition patterns on loss and identifies optimal epoch ranges for fixed datasets, such as approximately four epochs for 2B-parameter models. Ultimately, it provides both theoretical grounding and empirical guidance for configuring pre-training in data-limited settings.
📝 Abstract
As pretraining increasingly repeats data, every run faces three questions: how many epochs to take, how that number should change with model size, and whether anything besides the epoch count matters. We answer them by pricing a repeated token against two references: one epoch on the same data, which gives its value, and fresh data at equal compute, which gives its cost. Against fresh data, the cost of repetition follows a single variable, the number of extra epochs divided by the unique tokens per parameter. Against the same data, a second epoch is worth nearly as much as a fresh one, and repeated tokens fall to half the value of fresh ones after a critical epoch count that grows with the training budget per parameter but hardly with model size. With unique data fixed, the predicted compute-optimal run grows model size and epochs together until loss stops improving, near the critical epoch count. The same variable accounts for the direction of size trends that appear to conflict: larger models tolerate fewer epochs when the corpus is fixed, from about 15 at 127M to 4 at 2B parameters, but not when unique data grow with the model. Counts alone do not determine loss: at identical counts, replaying shards consecutively raises loss by up to 0.46~bits per byte, concentrating repeats on fewer samples also raises it, lower-entropy sources degrade faster with repetition, and re-tokenizing repeats helps only under heavy repetition. These results offer an empirical guide to pretraining when unique data, rather than compute, are the binding constraint.
Problem

Research questions and friction points this paper is trying to address.

multi-epoch pretraining
repeated tokens
scaling laws
compute-optimal training
data repetition
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-epoch pretraining
scaling geometry
repeated token value
compute-optimal training
data-constrained pretraining
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.