continual pre-training

Designs and implements workflows and algorithms to continue pretraining existing model weights on additional data streams or tasks (including specialization pretraining and abbreviation cpt), performing sequential adaptation of pretrained models to new domains, languages, or data while selecting data, scheduling updates, and tuning optimization. Builds and evaluates techniques to preserve prior knowledge (e.g., mitigate catastrophic forgetting) and to measure gains via downstream fine‑tuning or evaluation metrics compared to unrelated pretraining baselines.

continualpre-training

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$211K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Learning Dynamics in Continual Pre-Training for Large Language Models

May 12, 2025
XW
Xingjin Wang
🏛️ University of Chinese Academy of Sciences | Chinese Academy of Sciences | Ritzz-AI

This work investigates the learning dynamics of continual pretraining (CPT) for large language models, focusing on the co-evolution of general capabilities and downstream domain performance with training steps. Addressing the challenge that validation loss is analytically intractable due to coupled distributional shift and learning rate annealing—which impedes hyperparameter tuning—we propose, for the first time, a CPT scaling law that explicitly decouples these two effects, enabling accurate loss prediction across diverse learning rate schedules. Through theoretical modeling, loss dynamics analysis, and extensive experiments across multiple datasets and scheduling strategies, the law demonstrates strong empirical alignment under varied CPT configurations. It provides principled guidance for selecting critical hyperparameters—including peak learning rate and replay ratio—thereby establishing an interpretable, predictive optimization framework to balance model generality and domain-specific adaptability.

Customizing hyper-parameters for general vs domain-specific performanceModeling CPT loss curve transition and scaling lawsUnderstanding learning dynamics in continual pre-training for LLMs

ADEPT: Continual Pretraining via Adaptive Expansion and Dynamic Decoupled Tuning

Oct 11, 2025
JZ
Jinyang Zhang
🏛️ Peking University | Zhejiang University

To address catastrophic forgetting and limited domain capacity in large language models (LLMs) during continual pretraining (CPT), this paper proposes an Adaptive Expansion and Dynamic Decoupled Tuning framework. Methodologically, it introduces a novel functionality-aware hierarchical selective expansion mechanism, integrated with unit-level importance-aware decoupled optimization and asymmetric learning rate scheduling, enabling synergistic modeling of general capability retention and domain-specific knowledge injection. Its key innovation lies in functionally decoupling parameter expansion from parameter updating—thereby eliminating their entanglement. Experiments demonstrate that tuning only 15% of parameters reduces training time by over 50%, while outperforming full-parameter fine-tuning on mathematical and medical benchmarks: general capability improves by 5.76% and domain-specific performance by 5.58%.

Addresses catastrophic forgetting in continual pretraining of large language modelsSeparates general and domain learning via decoupled parameter optimizationSolves limited domain capacity through adaptive layer expansion strategies

Merging Continual Pretraining Models for Domain-Specialized LLMs: A Case Study in Finance

Nov 04, 2025
KU
Kentaro Ueda
🏛️ NARA Institute of Science and Technology | Université Grenoble Alpes

General-purpose large language models (LLMs) exhibit deficiencies in domain-specific knowledge (e.g., finance), mathematical reasoning, and multilingual capabilities. Method: We propose constructing a high-performance financial-domain LLM by fusing multiple domain-specific continual pretraining (CPT) expert models—avoiding costly and unstable end-to-end multi-skill training. Contribution/Results: This work presents the first systematic study on CPT model fusion, introducing a three-stage evaluation framework (knowledge recovery, skill complementarity, capability emergence) and benchmarking Task Arithmetic, TIES, and DARE-TIES on an 18-task financial evaluation suite. Fusion effectively restores general knowledge, improves overall performance, and induces emergent cross-domain capabilities. TIES demonstrates superior robustness, while Task Arithmetic achieves strong performance but is highly sensitive to hyperparameters. Our framework establishes principled, efficient pathways for building multi-competency domain LLMs from existing expert model assets.

Addressing knowledge loss and performance gaps in multi-skill model integrationEvaluating merging methods for emergent cross-domain capabilities in financeMerging domain-specific continual pretraining models for specialized financial LLMs

Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training

Jul 07, 2025
SL
Song Lai
🏛️ HKISI | City University of Hong Kong | Institute of Automation | University of Chinese Academy of Sciences

This work investigates catastrophic forgetting in continual post-training (CPT), specifically comparing supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT). We find that RFT inherently preserves knowledge, maintaining or even enhancing general capabilities (e.g., MMMU, MMLU-Pro) across multi-task continual learning. We attribute this advantage to implicit KL regularization emerging during policy optimization and propose a rollout-based instance filtering algorithm to improve RFT’s training stability and efficiency. To enable systematic evaluation, we introduce the first CPT benchmark tailored for multimodal tasks, integrating chain-of-thought reasoning and KL divergence analysis. Experiments on a seven-stage sequential task setup demonstrate that our RFT method matches the performance of full multi-task learning—without requiring memory replay or parameter isolation mechanisms.

Compares SFT and RFT impacts on knowledge retention in continual post-trainingExplores implicit regularization in RFT as key to mitigating forgettingInvestigates catastrophic forgetting in SFT versus knowledge preservation in RFT

Investigating Continual Pretraining in Large Language Models: Insights and Implications

Feb 27, 2024
ÇY
Çağatay Yıldız
🏛️ University of Tübingen | Cohere for AI Community | Cohere for AI

This study addresses the challenge of continual pretraining of large language models (LLMs) in dynamic knowledge environments, aiming to balance assimilation of new knowledge with retention of prior knowledge. To this end, we introduce the first benchmark specifically designed for evaluating continual pretraining under evolving data distributions, enabling systematic analysis of the interplay among model scale, semantic structure of domain sequences, and knowledge transfer/forgetting. We propose a novel cross-domain adaptive evaluation paradigm and uncover three key findings: (i) smaller models (<1.5B parameters) exhibit high sensitivity to both learning and forgetting; (ii) semantically ordered domain sequences foster specialization, whereas random sequences enhance generalization and cross-domain transfer; and (iii) larger models consistently achieve lower perplexity. Empirical results demonstrate that our continual pretraining paradigm significantly improves downstream task performance across the GPT-2 family, with particularly pronounced gains for smaller models.

Examines model size impact on learning and forgetting.Explores continual pretraining in large language models.Measures adaptability to changing pretraining data landscapes.

Latest Papers

What's happening recently
View more

This work addresses the fragmented landscape of post-training adaptation techniques, which suffer from inconsistent terminology and a lack of unified comparative or governance frameworks. To resolve this, the paper introduces the first six-dimensional taxonomy—spanning mechanism, objective, data requirements, persistence, structural scope, and model type—that systematically integrates mainstream approaches such as fine-tuning, retrieval augmentation, prompt engineering, model editing, and machine unlearning. This framework clarifies conceptual boundaries and reveals evolutionary and compositional relationships among methods. Beyond standardizing terminology, it enables standardized technical documentation, model change tracking, and AI governance analysis. The study further identifies critical challenges, including evaluation rigor, reproducibility, continual adaptation, multimodal alignment, and governance-aware workflows.

AI governancefoundation modelsmodel modification

Existing theoretical frameworks struggle to explain why larger-scale pre-trained models substantially reduce sample complexity on downstream tasks. This work proposes a novel theoretical framework—termed “caulking”—inspired by parameter-efficient fine-tuning methods such as adapters, low-rank adaptation, and partial fine-tuning. It establishes, for the first time, a provable relationship between the scale of pre-trained models and the sample complexity of downstream tasks. By rigorously linking stronger pre-training capabilities to reduced data requirements in transfer learning, this study not only addresses a critical gap in current theoretical understanding but also provides a solid foundation for empirically observed scaling laws, demonstrating that enhanced pre-training capacity can significantly decrease the number of samples needed for effective downstream adaptation.

downstream taskspre-trained modelssample complexity

This work addresses the critical challenge in continual learning of mitigating catastrophic forgetting during downstream fine-tuning while preserving capabilities acquired during upstream training. The authors propose treating “robustness to subsequent fine-tuning” as a first-class objective in upstream training and systematically investigate data scheduling strategies across a three-stage pipeline—pretraining, post-training, and downstream fine-tuning. Their key finding is that early exposure to post-training data during pretraining—termed “early data exposure”—consistently outperforms pure post-training or conventional mixing strategies, yielding superior trade-offs between upstream knowledge retention and downstream task performance across model scales from 135M to 1B parameters. This approach complements regularization techniques such as replay and Dropout and, under fixed compute budgets, reveals an optimal data allocation scheme.

catastrophic forgettingearly exposurefine-tuning

This study addresses the lack of systematic evaluation regarding the adaptability of existing general-purpose or code-oriented language models to non-code software engineering (SE) texts, such as issue reports and commit messages. Under strictly controlled computational and token budgets, it presents the first fair comparison between continual pre-training (CPT) and pre-training from scratch (PTS) in terms of their impact on domain adaptation and general language understanding capabilities for both encoder and decoder architectures trained on SE corpora. The results demonstrate that CPT yields limited and inconsistent domain-specific gains while largely preserving general capabilities, whereas PTS consistently degrades performance across both dimensions, showing competitiveness only for small models under high token budgets. These findings empirically establish that reusing existing models is substantially more effective than training from scratch, offering practical guidance for efficient adaptation of language models in SE contexts.

domain adaptationlanguage modelspre-training

This work addresses the long-standing gap in systematic research on pretraining—a phase that fundamentally determines a model’s capability ceiling—hampered by industrial opacity and academic compute constraints. Leveraging industrial-scale computational resources and full scientific autonomy, we propose Data Darwinism, a novel framework featuring an L0–L9 taxonomy of data processing depth, establishing data curation rigor as a critical dimension alongside data volume. Through two-stage adaptive curriculum learning, we train a 3B-parameter model on 8 trillion tokens and conduct over 200 ablation studies. Our findings reveal domain-specific saturation dynamics and a combinatorial balancing mechanism, demonstrating that principled data processing substantially enhances model performance while preventing catastrophic degradation. The entire pipeline is open-sourced to foster cumulative progress in the science of pretraining.

data processinglarge language modelspretraining

Hot Scholars

RN

Rodrigo Nogueira

Founder and CEO of Maritaca AI
Deep LearningNatural Language ProcessingInformation Retrieval
SR

Surangika Ranathunga

Senior Lecturer, School of Mathematical and Computational Sciences, Massey University, New Zealand
Natural Language ProcessingMachine LearningLarge Language Models
CX

Caiming Xiong

Salesforce Research
Machine LearningNLPComputer VisionMultimedia
GZ

Ge Zhang

M-A-P, Bytedance, University of Waterloo
Natural Language ProcessingMultimodal IntelligenceMusic ProcessingAgent