From Spectra to Joint Schedules in LLM Pre-training: 3+3(+2) Scaling-Law Regimes

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how learning rate and batch size schedules dynamically alter the power-law loss curves of LLM pretraining. Based on a linear random feature model, it employs Volterra integral equations to disentangle error propagation mechanisms, analytically characterizing the interplay between intrinsic time and noise injection under joint scheduling. Core contributions include deriving exact spectral conditions that determine whether power laws are preserved or broken, and establishing a "3+3(+2)" propagation regime alongside a phase-dependent computational rate framework. Online SGD analysis and nanoGPT experiments confirm that matched B/η paths are equivalent in intrinsic time. Furthermore, the proposed surrogate model accurately fits empirical data, enabling effective identification of the scaling law regimes governing LLMs.
📝 Abstract
Power-law learning curves are often treated as fixed properties of a model and its data, although learning-rate and batch-size schedules can change the observed loss. We study this dependence in noisy online SGD with linear random features. Conditional on the representation, an exact Volterra equation separates two response components: a forcing term that propagates unresolved target error and a memory kernel that propagates stochastic-error injections. We prove that either component follows a power law if and only if its cumulative weighted spectral mass has the corresponding low-spectrum scaling; individual eigenvalues and target coefficients need not obey coordinatewise power laws. Under a joint schedule, intrinsic time $T_t=\sum_{s<t}η_s$ controls optimization progress, while $r_t=B_t/η_t$ controls noise injection. Their interaction yields sharp conditions under which a schedule preserves, changes, or destroys the clean power law, together with a memory ceiling on noise reduction. The power-law random-feature model realizes this mechanism in $3+3(+2)$ propagation regimes with phase-dependent compute rates. Controlled nanoGPT experiments show that (1) learning-rate and batch-size schedules with matched $B/η$ paths are nearly equivalent in intrinsic time, (2) a forcing-memory surrogate accurately predicts loss across schedules, and (3) its fitted exponents across real-world datasets identify the regime of LLMs in $3+3(+2)$ map.
Problem

Research questions and friction points this paper is trying to address.

scaling laws
LLM pre-training
learning rate schedule
batch size schedule
power-law learning curves
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scaling Laws
Volterra Equation
Joint Schedules
Random Features
LLM Pre-training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yichen Wang
Department of Computer Sciences, University of Wisconsin–Madison, USA
Fanghui Liu
Fanghui Liu
Assistant Professor, University of Warwick
Foundations of Modern ML
Y
Yudong Chen
Department of Computer Sciences, University of Wisconsin–Madison, USA