Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过提出终端收缩平均法(TSA),在大语言模型预训练中平衡优化进展与模型变异性,改进了学习率调度和估计器选择。
📝 Abstract
Large language model (LLM) pretraining conventionally returns the raw final iterate. This couples two design choices: the learning-rate schedule that generates the parameter trajectory and the estimator that constructs the deployed model (e.g. the raw final iterate or a checkpoint average). A schedule that promotes optimization progress may differ from one that minimizes variation in the raw final iterate. Separating these choices creates an opportunity to maintain progress late in training while reducing variation in the returned model. To this end, we propose \emph{Terminal Shrinkage Averaging (TSA)}, which interpolates between the raw final iterate and the average of recent checkpoints to balance recent progress against terminal variation. We analyze how TSA changes the preferred terminal learning-rate schedule under a local quadratic approximation and test this interaction through a sequence of controlled NanoChat experiments. Finally, we demonstrate that the resulting gains transfer to depth-22 NanoChat, where the combined schedule and estimator improve validation quality. A qualifying time-to-GPT-2 run also finishes faster than the public baseline used in our experiments, providing preliminary evidence of benchmark acceleration.
Problem

Research questions and friction points this paper is trying to address.

Large Language Model
Pretraining
Learning-rate Schedule
Estimator
Model Variation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Terminal Shrinkage Averaging
learning-rate schedule
estimator
optimization progress
variation reduction
💼 Related Jobs
No related jobs found.
A
Adam Ousherovitch
Department of Statistics, University of Michigan
Yixin Wang
Yixin Wang
University of Michigan
Bayesian statisticsMachine Learning