Score
Designs and implements workflows and algorithms to continue pretraining existing model weights on additional data streams or tasks (including specialization pretraining and abbreviation cpt), performing sequential adaptation of pretrained models to new domains, languages, or data while selecting data, scheduling updates, and tuning optimization. Builds and evaluates techniques to preserve prior knowledge (e.g., mitigate catastrophic forgetting) and to measure gains via downstream fine‑tuning or evaluation metrics compared to unrelated pretraining baselines.
This work investigates the learning dynamics of continual pretraining (CPT) for large language models, focusing on the co-evolution of general capabilities and downstream domain performance with training steps. Addressing the challenge that validation loss is analytically intractable due to coupled distributional shift and learning rate annealing—which impedes hyperparameter tuning—we propose, for the first time, a CPT scaling law that explicitly decouples these two effects, enabling accurate loss prediction across diverse learning rate schedules. Through theoretical modeling, loss dynamics analysis, and extensive experiments across multiple datasets and scheduling strategies, the law demonstrates strong empirical alignment under varied CPT configurations. It provides principled guidance for selecting critical hyperparameters—including peak learning rate and replay ratio—thereby establishing an interpretable, predictive optimization framework to balance model generality and domain-specific adaptability.
To address catastrophic forgetting and limited domain capacity in large language models (LLMs) during continual pretraining (CPT), this paper proposes an Adaptive Expansion and Dynamic Decoupled Tuning framework. Methodologically, it introduces a novel functionality-aware hierarchical selective expansion mechanism, integrated with unit-level importance-aware decoupled optimization and asymmetric learning rate scheduling, enabling synergistic modeling of general capability retention and domain-specific knowledge injection. Its key innovation lies in functionally decoupling parameter expansion from parameter updating—thereby eliminating their entanglement. Experiments demonstrate that tuning only 15% of parameters reduces training time by over 50%, while outperforming full-parameter fine-tuning on mathematical and medical benchmarks: general capability improves by 5.76% and domain-specific performance by 5.58%.
General-purpose large language models (LLMs) exhibit deficiencies in domain-specific knowledge (e.g., finance), mathematical reasoning, and multilingual capabilities. Method: We propose constructing a high-performance financial-domain LLM by fusing multiple domain-specific continual pretraining (CPT) expert models—avoiding costly and unstable end-to-end multi-skill training. Contribution/Results: This work presents the first systematic study on CPT model fusion, introducing a three-stage evaluation framework (knowledge recovery, skill complementarity, capability emergence) and benchmarking Task Arithmetic, TIES, and DARE-TIES on an 18-task financial evaluation suite. Fusion effectively restores general knowledge, improves overall performance, and induces emergent cross-domain capabilities. TIES demonstrates superior robustness, while Task Arithmetic achieves strong performance but is highly sensitive to hyperparameters. Our framework establishes principled, efficient pathways for building multi-competency domain LLMs from existing expert model assets.
This work investigates catastrophic forgetting in continual post-training (CPT), specifically comparing supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT). We find that RFT inherently preserves knowledge, maintaining or even enhancing general capabilities (e.g., MMMU, MMLU-Pro) across multi-task continual learning. We attribute this advantage to implicit KL regularization emerging during policy optimization and propose a rollout-based instance filtering algorithm to improve RFT’s training stability and efficiency. To enable systematic evaluation, we introduce the first CPT benchmark tailored for multimodal tasks, integrating chain-of-thought reasoning and KL divergence analysis. Experiments on a seven-stage sequential task setup demonstrate that our RFT method matches the performance of full multi-task learning—without requiring memory replay or parameter isolation mechanisms.
This study addresses the challenge of continual pretraining of large language models (LLMs) in dynamic knowledge environments, aiming to balance assimilation of new knowledge with retention of prior knowledge. To this end, we introduce the first benchmark specifically designed for evaluating continual pretraining under evolving data distributions, enabling systematic analysis of the interplay among model scale, semantic structure of domain sequences, and knowledge transfer/forgetting. We propose a novel cross-domain adaptive evaluation paradigm and uncover three key findings: (i) smaller models (<1.5B parameters) exhibit high sensitivity to both learning and forgetting; (ii) semantically ordered domain sequences foster specialization, whereas random sequences enhance generalization and cross-domain transfer; and (iii) larger models consistently achieve lower perplexity. Empirical results demonstrate that our continual pretraining paradigm significantly improves downstream task performance across the GPT-2 family, with particularly pronounced gains for smaller models.
This work addresses the fragmented landscape of post-training adaptation techniques, which suffer from inconsistent terminology and a lack of unified comparative or governance frameworks. To resolve this, the paper introduces the first six-dimensional taxonomy—spanning mechanism, objective, data requirements, persistence, structural scope, and model type—that systematically integrates mainstream approaches such as fine-tuning, retrieval augmentation, prompt engineering, model editing, and machine unlearning. This framework clarifies conceptual boundaries and reveals evolutionary and compositional relationships among methods. Beyond standardizing terminology, it enables standardized technical documentation, model change tracking, and AI governance analysis. The study further identifies critical challenges, including evaluation rigor, reproducibility, continual adaptation, multimodal alignment, and governance-aware workflows.
Existing theoretical frameworks struggle to explain why larger-scale pre-trained models substantially reduce sample complexity on downstream tasks. This work proposes a novel theoretical framework—termed “caulking”—inspired by parameter-efficient fine-tuning methods such as adapters, low-rank adaptation, and partial fine-tuning. It establishes, for the first time, a provable relationship between the scale of pre-trained models and the sample complexity of downstream tasks. By rigorously linking stronger pre-training capabilities to reduced data requirements in transfer learning, this study not only addresses a critical gap in current theoretical understanding but also provides a solid foundation for empirically observed scaling laws, demonstrating that enhanced pre-training capacity can significantly decrease the number of samples needed for effective downstream adaptation.
This work addresses the critical challenge in continual learning of mitigating catastrophic forgetting during downstream fine-tuning while preserving capabilities acquired during upstream training. The authors propose treating “robustness to subsequent fine-tuning” as a first-class objective in upstream training and systematically investigate data scheduling strategies across a three-stage pipeline—pretraining, post-training, and downstream fine-tuning. Their key finding is that early exposure to post-training data during pretraining—termed “early data exposure”—consistently outperforms pure post-training or conventional mixing strategies, yielding superior trade-offs between upstream knowledge retention and downstream task performance across model scales from 135M to 1B parameters. This approach complements regularization techniques such as replay and Dropout and, under fixed compute budgets, reveals an optimal data allocation scheme.
This study addresses the lack of systematic evaluation regarding the adaptability of existing general-purpose or code-oriented language models to non-code software engineering (SE) texts, such as issue reports and commit messages. Under strictly controlled computational and token budgets, it presents the first fair comparison between continual pre-training (CPT) and pre-training from scratch (PTS) in terms of their impact on domain adaptation and general language understanding capabilities for both encoder and decoder architectures trained on SE corpora. The results demonstrate that CPT yields limited and inconsistent domain-specific gains while largely preserving general capabilities, whereas PTS consistently degrades performance across both dimensions, showing competitiveness only for small models under high token budgets. These findings empirically establish that reusing existing models is substantially more effective than training from scratch, offering practical guidance for efficient adaptation of language models in SE contexts.
This work addresses the long-standing gap in systematic research on pretraining—a phase that fundamentally determines a model’s capability ceiling—hampered by industrial opacity and academic compute constraints. Leveraging industrial-scale computational resources and full scientific autonomy, we propose Data Darwinism, a novel framework featuring an L0–L9 taxonomy of data processing depth, establishing data curation rigor as a critical dimension alongside data volume. Through two-stage adaptive curriculum learning, we train a 3B-parameter model on 8 trillion tokens and conduct over 200 ablation studies. Our findings reveal domain-specific saturation dynamics and a combinatorial balancing mechanism, demonstrating that principled data processing substantially enhances model performance while preventing catastrophic degradation. The entire pipeline is open-sourced to foster cumulative progress in the science of pretraining.