Score
Designs and executes continued pretraining of pretrained language models on corpora selected to match a target domain, including constructing domain-specific corpora, tokenization/vocabulary changes, and initialization strategies. Measures and analyzes how this additional pretraining affects downstream task performance, transferability, calibration and out-of-distribution errors, and trade-offs introduced by vocabulary or formatting changes.
The “mid-training” phase—occurring between pretraining and post-training—has been largely overlooked in large language model (LLM) development, despite its critical role in balancing targeted capability enhancement with preservation of foundational language modeling performance. Method: We formally define and categorize mid-training for the first time, proposing a multi-stage optimization framework encompassing data curation, curriculum learning, continued pretraining, instruction tuning, and architectural expansion. Contribution/Results: Empirical evaluation demonstrates that our approach systematically improves target capabilities—including mathematical reasoning, code generation, complex reasoning, and long-context understanding—while robustly maintaining general language modeling performance. This work establishes the first theoretical framework and practical guidelines for mid-training, enabling reproducible, controllable, and efficient LLM capability evolution and domain-specific customization.
Pretraining and instruction tuning exhibit syntactic and task-distribution mismatches, leading to catastrophic forgetting of domain-specific knowledge—particularly in mathematics and code. Method: We introduce high-quality instruction data during the late pretraining phase (“midtraining”) and conduct controlled ablation studies on models trained from scratch, using diverse supervised fine-tuning datasets. Contribution/Results: We provide the first empirical evidence that midtraining functions as an effective domain adaptation technique, substantially mitigating knowledge forgetting in mathematical and programming domains. Its efficacy depends primarily on the timing of intervention—not on the proportion of instruction data mixed into pretraining. Under equal data budgets, midtraining achieves significantly lower domain-specific validation loss compared to continued pretraining. Our findings deliver causal, stage-level insights into training dynamics, establishing midtraining as a principled strategy for aligning pretraining with downstream task distributions.
This study addresses the challenge of continual pretraining of large language models (LLMs) in dynamic knowledge environments, aiming to balance assimilation of new knowledge with retention of prior knowledge. To this end, we introduce the first benchmark specifically designed for evaluating continual pretraining under evolving data distributions, enabling systematic analysis of the interplay among model scale, semantic structure of domain sequences, and knowledge transfer/forgetting. We propose a novel cross-domain adaptive evaluation paradigm and uncover three key findings: (i) smaller models (<1.5B parameters) exhibit high sensitivity to both learning and forgetting; (ii) semantically ordered domain sequences foster specialization, whereas random sequences enhance generalization and cross-domain transfer; and (iii) larger models consistently achieve lower perplexity. Empirical results demonstrate that our continual pretraining paradigm significantly improves downstream task performance across the GPT-2 family, with particularly pronounced gains for smaller models.
This study investigates the causal impact of code–natural language mixed pretraining on large language model performance. We systematically vary the code proportion—under both additive and competitive data mixing regimes—while maintaining a uniform Transformer architecture for pretraining, and evaluate models across diverse benchmarks including BigBench, semantic parsing, syntactic transformation, and commonsense reasoning. Our work establishes, for the first time, a causal relationship between code pretraining ratio and downstream task performance. We find that higher code proportions significantly enhance structured reasoning capabilities (e.g., semantic parsing and mathematical reasoning) but degrade sensitivity to linguistic structure (syntax and morphology) and impair commonsense reasoning. These results reveal a task-selective gain mechanism induced by code pretraining, wherein structural inductive biases from code benefit formal reasoning at the cost of natural language understanding. The findings provide both theoretical grounding and empirical evidence for principled, capability-aware pretraining data composition.
This study investigates the intrinsic synergy and trade-offs between pretraining and fine-tuning in large language models (LLMs). Methodologically, we propose a multi-stage fine-tuning analysis framework leveraging intermediate pretraining checkpoints, systematically evaluating capability improvement, adaptation to new knowledge, retention of prior knowledge, and prompt robustness across 18 diverse datasets. Key findings are: (1) continued pretraining implicitly enhances downstream fine-tuning performance; (2) fine-tuning yields substantial gains on weak-task capabilities but induces domain-specific knowledge forgetting; (3) fine-tuning exacerbates prompt sensitivity, whereas additional pretraining effectively mitigates this effect. Crucially, we quantitatively demonstrate the reversibility of both knowledge forgetting and prompt sensitivity—establishing that pretraining quality fundamentally bounds fine-tuning efficacy. Our work provides the first reproducible empirical guidelines and standardized evaluation protocols for optimizing the pretraining–fine-tuning pipeline.
This work investigates the effectiveness and underlying mechanisms of continual pretraining (CPT) in generative unsupervised domain adaptation (UDA). Addressing the gap that existing UDA research focuses predominantly on discriminative methods while generative UDA remains underexplored, we present the first systematic evaluation of CPT for generative UDA. We propose a CPT paradigm grounded in masked language modeling (MLM), integrated with domain-invariant representation learning. Through extensive ablation studies across diverse model architectures, fine-tuning strategies, and data scales, we demonstrate its robust generalizability. Results show that CPT substantially improves target-domain classification accuracy. Its core mechanism lies in implicitly acquiring downstream classification capability by predicting task-informative masked tokens during MLM. Moreover, we theoretically establish consistency between CPT and instruction tuning in terms of task-guided representation learning, revealing a unified principle for effective adaptation.
High experimental costs and difficulties in conducting controlled, multi-condition studies hinder pretraining research for large language models (LLMs). To address this, we propose a “single-training, multiple-experiments” paradigm: ten heterogeneous experiments—including knowledge acquisition, mathematical reasoning, and others—are executed in parallel during a single 1.5B-parameter LLM pretraining run. Leveraging controlled-variable design, dynamic data injection, interactive detection, and contamination analysis, we ensure negligible cross-experiment interference. This approach dramatically improves research efficiency—reproducing established findings and enabling novel explorations—while incurring virtually no additional computational overhead or performance degradation, achieving up to 90% compute savings. Our core contribution is the first systematic realization of a scientific experimentation framework for LLM pretraining that supports concurrent multi-task learning, multi-hypothesis testing, and full reproducibility.
Pre-trained tokenizers suffer from inefficient vocabulary expansion and imprecise pruning of redundant tokens during cross-domain or cross-lingual transfer. Method: This paper proposes a dynamic vocabulary optimization framework based on continued Byte-Pair Encoding (BPE) training. It incrementally integrates new vocabulary by extending the original BPE merge process, thereby improving token utilization; additionally, it introduces, for the first time, a leaf-node pruning strategy grounded in the BPE tree structure, enabling controllable and interpretable vocabulary reduction without compromising model performance. Contribution/Results: Experiments across multilingual settings and model families (e.g., BERT, XLM-R) demonstrate an average 15% vocabulary compression, substantial improvements in tokenization efficiency, and a 32% increase in usage rate of newly added tokens—establishing a robust, efficient paradigm for tokenizer customization and adaptation.
This study addresses the lack of systematic evaluation regarding the adaptability of existing general-purpose or code-oriented language models to non-code software engineering (SE) texts, such as issue reports and commit messages. Under strictly controlled computational and token budgets, it presents the first fair comparison between continual pre-training (CPT) and pre-training from scratch (PTS) in terms of their impact on domain adaptation and general language understanding capabilities for both encoder and decoder architectures trained on SE corpora. The results demonstrate that CPT yields limited and inconsistent domain-specific gains while largely preserving general capabilities, whereas PTS consistently degrades performance across both dimensions, showing competitiveness only for small models under high token budgets. These findings empirically establish that reusing existing models is substantially more effective than training from scratch, offering practical guidance for efficient adaptation of language models in SE contexts.
To address subword segmentation redundancy, excessive sequence length, and inference latency in large language models (LLMs) when processing out-of-domain text—caused by vocabulary mismatch—this paper proposes a **length-preserving vocabulary expansion method**. Leveraging domain-specific term frequency analysis, it seamlessly injects high-frequency domain tokens into the pretrained tokenizer and introduces an optimization algorithm guaranteeing that the tokenized sequence length never exceeds that produced by the original vocabulary. The method requires no model retraining, preserving tokenizer efficiency and backward compatibility. Evaluated on real-world e-commerce data, it reduces input sequence length by up to 20%, significantly lowering inference latency while maintaining zero performance degradation on downstream tasks. Its core contribution is the first realization of **strictly length-constrained, domain-adaptive vocabulary expansion**, uniquely balancing computational efficiency, system compatibility, and practical deployability.
This work addresses the challenge of optimally allocating computational resources between general pretraining and domain-specific fine-tuning for language models in multi-domain scenarios. The authors propose a scaling-law-based optimization method that trains multiple models in parallel on general corpora and leverages empirical scaling laws to accurately predict loss across varying model sizes and data volumes, enabling reliable extrapolation to larger scales. This approach dynamically determines the optimal split of compute between general pretraining and continued domain-adaptive pretraining. Experimental results demonstrate consistent and significant performance gains across diverse model scales and computational budgets on commonsense and reasoning benchmarks, marking the first achievement of efficient, cross-scale and cross-domain resource allocation for large language models.