Learning Dynamics in Continual Pre-Training for Large Language Models

📅 2025-05-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the learning dynamics of continual pretraining (CPT) for large language models, focusing on the co-evolution of general capabilities and downstream domain performance with training steps. Addressing the challenge that validation loss is analytically intractable due to coupled distributional shift and learning rate annealing—which impedes hyperparameter tuning—we propose, for the first time, a CPT scaling law that explicitly decouples these two effects, enabling accurate loss prediction across diverse learning rate schedules. Through theoretical modeling, loss dynamics analysis, and extensive experiments across multiple datasets and scheduling strategies, the law demonstrates strong empirical alignment under varied CPT configurations. It provides principled guidance for selecting critical hyperparameters—including peak learning rate and replay ratio—thereby establishing an interpretable, predictive optimization framework to balance model generality and domain-specific adaptability.

Technology Category

Natural Language Processing: Learning & Optimization for NLPMachine Learning: Life-Long and Continual LearningSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingWeb Mining and Content Analysis: Large pretrained models with web dataUser Modeling, Personalization and Recommendation: Practical large-scale studies of user experience
📝 Abstract
Continual Pre-Training (CPT) has become a popular and effective method to apply strong foundation models to specific downstream tasks. In this work, we explore the learning dynamics throughout the CPT process for large language models. We specifically focus on how general and downstream domain performance evolves at each training step, with domain performance measured via validation losses. We have observed that the CPT loss curve fundamentally characterizes the transition from one curve to another hidden curve, and could be described by decoupling the effects of distribution shift and learning rate annealing. We derive a CPT scaling law that combines the two factors, enabling the prediction of loss at any (continual) training steps and across learning rate schedules (LRS) in CPT. Our formulation presents a comprehensive understanding of several critical factors in CPT, including loss potential, peak learning rate, training steps, replay ratio, etc. Moreover, our approach can be adapted to customize training hyper-parameters to different CPT goals such as balancing general and domain-specific performance. Extensive experiments demonstrate that our scaling law holds across various CPT datasets and training hyper-parameters.
Problem

Research questions and friction points this paper is trying to address.

Understanding learning dynamics in continual pre-training for LLMs
Modeling CPT loss curve transition and scaling laws
Customizing hyper-parameters for general vs domain-specific performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decoupling distribution shift and learning rate effects
Deriving CPT scaling law for loss prediction
Customizing hyper-parameters for CPT goals
X
Xingjin Wang
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China; State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China
H
Howe Tissue
L
Lu Wang
Ritzz-AI
L
Linjing Li
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China; State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China
D
Daniel Dajun Zeng
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China; State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China