domain-adaptive pretraining

Designs and executes continued pretraining of pretrained language models on corpora selected to match a target domain, including constructing domain-specific corpora, tokenization/vocabulary changes, and initialization strategies. Measures and analyzes how this additional pretraining affects downstream task performance, transferability, calibration and out-of-distribution errors, and trade-offs introduced by vocabulary or formatting changes.

domain-adaptivepretraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.58
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Midtraining Bridges Pretraining and Posttraining Distributions

Oct 16, 2025
EL
Emmy Liu
🏛️ Carnegie Mellon University

Pretraining and instruction tuning exhibit syntactic and task-distribution mismatches, leading to catastrophic forgetting of domain-specific knowledge—particularly in mathematics and code. Method: We introduce high-quality instruction data during the late pretraining phase (“midtraining”) and conduct controlled ablation studies on models trained from scratch, using diverse supervised fine-tuning datasets. Contribution/Results: We provide the first empirical evidence that midtraining functions as an effective domain adaptation technique, substantially mitigating knowledge forgetting in mathematical and programming domains. Its efficacy depends primarily on the timing of intervention—not on the proportion of instruction data mixed into pretraining. Under equal data budgets, midtraining achieves significantly lower domain-specific validation loss compared to continued pretraining. Our findings deliver causal, stage-level insights into training dynamics, establishing midtraining as a principled strategy for aligning pretraining with downstream task distributions.

It reduces syntactic disparities between pretraining and fine-tuning domainsMidtraining bridges pretraining and posttraining data distribution gapsMidtraining outperforms continued pretraining by minimizing catastrophic forgetting

Investigating Continual Pretraining in Large Language Models: Insights and Implications

Feb 27, 2024
ÇY
Çağatay Yıldız
🏛️ University of Tübingen | Cohere for AI Community | Cohere for AI

This study addresses the challenge of continual pretraining of large language models (LLMs) in dynamic knowledge environments, aiming to balance assimilation of new knowledge with retention of prior knowledge. To this end, we introduce the first benchmark specifically designed for evaluating continual pretraining under evolving data distributions, enabling systematic analysis of the interplay among model scale, semantic structure of domain sequences, and knowledge transfer/forgetting. We propose a novel cross-domain adaptive evaluation paradigm and uncover three key findings: (i) smaller models (<1.5B parameters) exhibit high sensitivity to both learning and forgetting; (ii) semantically ordered domain sequences foster specialization, whereas random sequences enhance generalization and cross-domain transfer; and (iii) larger models consistently achieve lower perplexity. Empirical results demonstrate that our continual pretraining paradigm significantly improves downstream task performance across the GPT-2 family, with particularly pronounced gains for smaller models.

Examines model size impact on learning and forgetting.Explores continual pretraining in large language models.Measures adaptability to changing pretraining data landscapes.

How Does Code Pretraining Affect Language Model Task Performance?

Sep 06, 2024
JP
Jackson Petty
🏛️ New York University | Google

This study investigates the causal impact of code–natural language mixed pretraining on large language model performance. We systematically vary the code proportion—under both additive and competitive data mixing regimes—while maintaining a uniform Transformer architecture for pretraining, and evaluate models across diverse benchmarks including BigBench, semantic parsing, syntactic transformation, and commonsense reasoning. Our work establishes, for the first time, a causal relationship between code pretraining ratio and downstream task performance. We find that higher code proportions significantly enhance structured reasoning capabilities (e.g., semantic parsing and mathematical reasoning) but degrade sensitivity to linguistic structure (syntax and morphology) and impair commonsense reasoning. These results reveal a task-selective gain mechanism induced by code pretraining, wherein structural inductive biases from code benefit formal reasoning at the cost of natural language understanding. The findings provide both theoretical grounding and empirical evidence for principled, capability-aware pretraining data composition.

Causal connection between code and language dataEffect of code pretraining on language modelsImpact of code mixture on task performance

This study investigates the intrinsic synergy and trade-offs between pretraining and fine-tuning in large language models (LLMs). Methodologically, we propose a multi-stage fine-tuning analysis framework leveraging intermediate pretraining checkpoints, systematically evaluating capability improvement, adaptation to new knowledge, retention of prior knowledge, and prompt robustness across 18 diverse datasets. Key findings are: (1) continued pretraining implicitly enhances downstream fine-tuning performance; (2) fine-tuning yields substantial gains on weak-task capabilities but induces domain-specific knowledge forgetting; (3) fine-tuning exacerbates prompt sensitivity, whereas additional pretraining effectively mitigates this effect. Crucially, we quantitatively demonstrate the reversibility of both knowledge forgetting and prompt sensitivity—establishing that pretraining quality fundamentally bounds fine-tuning efficacy. Our work provides the first reproducible empirical guidelines and standardized evaluation protocols for optimizing the pretraining–fine-tuning pipeline.

Fine-tuningLarge Language ModelsTransfer Learning

How Useful is Continued Pre-Training for Generative Unsupervised Domain Adaptation?

Jan 31, 2024
RU
Rheeya Uppaal
🏛️ University of Wisconsin-Madison

This work investigates the effectiveness and underlying mechanisms of continual pretraining (CPT) in generative unsupervised domain adaptation (UDA). Addressing the gap that existing UDA research focuses predominantly on discriminative methods while generative UDA remains underexplored, we present the first systematic evaluation of CPT for generative UDA. We propose a CPT paradigm grounded in masked language modeling (MLM), integrated with domain-invariant representation learning. Through extensive ablation studies across diverse model architectures, fine-tuning strategies, and data scales, we demonstrate its robust generalizability. Results show that CPT substantially improves target-domain classification accuracy. Its core mechanism lies in implicitly acquiring downstream classification capability by predicting task-informative masked tokens during MLM. Moreover, we theoretically establish consistency between CPT and instruction tuning in terms of task-guided representation learning, revealing a unified principle for effective adaptation.

Evaluates Continued Pre-Training for generative Unsupervised Domain Adaptation.Explores trade-offs between CPT and domain invariance methods.Investigates CPT's impact on classification in unlabeled target domains.

Latest Papers

What's happening recently
View more

Train Once, Answer All: Many Pretraining Experiments for the Cost of One

Sep 27, 2025
SB
Sebastian Bordt
🏛️ University of Töbingen | Töbingen AI Center | Independent Researcher

High experimental costs and difficulties in conducting controlled, multi-condition studies hinder pretraining research for large language models (LLMs). To address this, we propose a “single-training, multiple-experiments” paradigm: ten heterogeneous experiments—including knowledge acquisition, mathematical reasoning, and others—are executed in parallel during a single 1.5B-parameter LLM pretraining run. Leveraging controlled-variable design, dynamic data injection, interactive detection, and contamination analysis, we ensure negligible cross-experiment interference. This approach dramatically improves research efficiency—reproducing established findings and enabling novel explorations—while incurring virtually no additional computational overhead or performance degradation, achieving up to 90% compute savings. Our core contribution is the first systematic realization of a scientific experimentation framework for LLM pretraining that supports concurrent multi-task learning, multi-hypothesis testing, and full reproducibility.

Enabling rigorous scientific research with limited compute budgetReducing computational cost of multiple pretraining experimentsSimultaneously conducting diverse experiments in single training run

Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pre-trained Models

Dec 03, 2025
TP
Taido Purason
🏛️ University of Tartu | Technical University of Applied Sciences Würzburg-Schweinfurt

Pre-trained tokenizers suffer from inefficient vocabulary expansion and imprecise pruning of redundant tokens during cross-domain or cross-lingual transfer. Method: This paper proposes a dynamic vocabulary optimization framework based on continued Byte-Pair Encoding (BPE) training. It incrementally integrates new vocabulary by extending the original BPE merge process, thereby improving token utilization; additionally, it introduces, for the first time, a leaf-node pruning strategy grounded in the BPE tree structure, enabling controllable and interpretable vocabulary reduction without compromising model performance. Contribution/Results: Experiments across multilingual settings and model families (e.g., BERT, XLM-R) demonstrate an average 15% vocabulary compression, substantial improvements in tokenization efficiency, and a 32% increase in usage rate of newly added tokens—establishing a robust, efficient paradigm for tokenizer customization and adaptation.

Adapt tokenizers for new domains or languages efficientlyExtend vocabulary without creating unused tokensPrune redundant tokens while maintaining model quality

This study addresses the lack of systematic evaluation regarding the adaptability of existing general-purpose or code-oriented language models to non-code software engineering (SE) texts, such as issue reports and commit messages. Under strictly controlled computational and token budgets, it presents the first fair comparison between continual pre-training (CPT) and pre-training from scratch (PTS) in terms of their impact on domain adaptation and general language understanding capabilities for both encoder and decoder architectures trained on SE corpora. The results demonstrate that CPT yields limited and inconsistent domain-specific gains while largely preserving general capabilities, whereas PTS consistently degrades performance across both dimensions, showing competitiveness only for small models under high token budgets. These findings empirically establish that reusing existing models is substantially more effective than training from scratch, offering practical guidance for efficient adaptation of language models in SE contexts.

domain adaptationlanguage modelspre-training

Vocabulary Customization for Efficient Domain-Specific LLM Deployment

Sep 30, 2025
CH
Christian Herold
🏛️ eBay Inc.

To address subword segmentation redundancy, excessive sequence length, and inference latency in large language models (LLMs) when processing out-of-domain text—caused by vocabulary mismatch—this paper proposes a **length-preserving vocabulary expansion method**. Leveraging domain-specific term frequency analysis, it seamlessly injects high-frequency domain tokens into the pretrained tokenizer and introduces an optimization algorithm guaranteeing that the tokenized sequence length never exceeds that produced by the original vocabulary. The method requires no model retraining, preserving tokenizer efficiency and backward compatibility. Evaluated on real-world e-commerce data, it reduces input sequence length by up to 20%, significantly lowering inference latency while maintaining zero performance degradation on downstream tasks. Its core contribution is the first realization of **strictly length-constrained, domain-adaptive vocabulary expansion**, uniquely balancing computational efficiency, system compatibility, and practical deployability.

Addresses vocabulary mismatch in LLMs for domain-specific text processingMaintains predictive quality while shortening input sequences by up to 20%Reduces token fertility and improves inference speed through vocabulary augmentation

This work addresses the challenge of optimally allocating computational resources between general pretraining and domain-specific fine-tuning for language models in multi-domain scenarios. The authors propose a scaling-law-based optimization method that trains multiple models in parallel on general corpora and leverages empirical scaling laws to accurately predict loss across varying model sizes and data volumes, enabling reliable extrapolation to larger scales. This approach dynamically determines the optimal split of compute between general pretraining and continued domain-adaptive pretraining. Experimental results demonstrate consistent and significant performance gains across diverse model scales and computational budgets on commonsense and reasoning benchmarks, marking the first achievement of efficient, cross-scale and cross-domain resource allocation for large language models.

compute allocationlanguage modelsmulti-domain

Hot Scholars

DR

Daniel Rueckert

Technical University of Munich and Imperial College London
Machine LearningMedical Image ComputingBiomedical Image AnalysisComputer Vision
YK

Yuta Koreeda

Hitachi, Ltd., Hitachi America, Ltd., Stanford CS
natural language processingmachine learningrobotcomputer assisted surgery
AA

Areej Alhothali

Associate Professor of Computer Science, King Abulaziz University
Machine learningNatural language processingAffective ComputingSentiment analysis
PS

Philip S. Yu

Professor of Computer Science, University of Illinons at Chicago
Data miningDatabasePrivacy
SF

Soukaina Filali Boubrahimi

Associate Professor of Computer Science, Utah State University, Logan, Utah, USA
Time Series Data MiningDatabase SystemsData MiningSpatio-Temporal Pattern Mining