design pretraining strategies

Designs and evaluates pretraining regimes for machine learning models, specifying pretraining objectives, optimization and data-selection methods, curricula and schedules for continued/continual/continuous pretraining, and techniques for efficient pretraining. Also develops procedures and trade-offs for combining pretraining with subsequent fine-tuning to meet target performance, compute, or data-efficiency constraints.

designpretrainingstrategies

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.78
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$217K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Beyond algorithm hyperparameters: on preprocessing hyperparameters and associated pitfalls in machine learning applications

Dec 04, 2024
CS
Christina Sauer
🏛️ LMU Munich | Munich Center for Machine Learning | Medical University of Vienna

This paper identifies a systemic issue in machine learning: preprocessing hyperparameters—such as missing-value imputation strategies—are frequently overlooked yet substantially bias model evaluation. Current practice often involves informal, post-hoc tuning of preprocessing steps, leading to optimistic performance estimates and irreproducible results. To address this, the authors formally distinguish and empirically analyze the coupling effects between algorithmic and preprocessing hyperparameters. Using a modular supervised learning workflow model, controlled variable experiments, replication of canonical case studies, and bias diagnostics, they quantify the resulting optimistic bias. Key contributions include: (1) establishing preprocessing hyperparameters as equally critical as algorithmic ones; (2) proposing formal modeling principles to eliminate informal preprocessing tuning; and (3) delivering actionable reporting guidelines for ML practitioners, thereby significantly enhancing model credibility and reproducibility.

Addresses overlooked preprocessing hyperparameters in ML model tuningAims to improve predictive modeling quality and reportingHighlights pitfalls in informal preprocessing optimization practices

Parameter-Efficient Continual Fine-Tuning: A Survey

Apr 18, 2025
EN
Eric Nuertey Coleman
🏛️ University of Pisa | Politecnico di Torino | University of Auckland | Indian Institute of Technology | University of Warwick

This paper addresses the fundamental challenge of balancing catastrophic forgetting and parameter efficiency when large pre-trained models continuously adapt to dynamic task streams. To this end, we propose the first unified theoretical framework for Parameter-Efficient Continual Fine-Tuning (PECFT). Our framework systematically organizes existing approaches along three dimensions: method taxonomy, evaluation metrics, and core challenges—integrating Parameter-Efficient Fine-Tuning (PEFT) techniques (e.g., adapters, LoRA, prompt tuning) with continual learning strategies (e.g., replay, regularization, architecture expansion). Through a comprehensive review of over 100 studies, we identify key trade-offs between performance and efficiency, and pinpoint scalable memory mechanisms and task-aware parameter updates as critical research frontiers. This work bridges a significant gap at the intersection of continual learning and PEFT, providing both theoretical foundations and practical guidelines for efficient, sustainable adaptation of large language models.

Addressing catastrophic forgetting in continual learning scenariosEnhancing parameter-efficient fine-tuning for dynamic environmentsSurveying methods for lifelong adaptation of large pre-trained models

In large-scale pretraining, learning rate scheduling critically influences both training efficiency and model performance. This work proposes two paradigms—Fitting and Transfer. The Fitting paradigm establishes, for the first time, a scaling law for learning rate search factors, reducing hyperparameter tuning complexity from O(n³) to O(n·C_D·C_η). The Transfer paradigm extends μTransfer to Mixture-of-Experts (MoE) architectures and generalizes it across multiple hyperparameter dimensions, including depth, weight decay, and token length. Empirical results demonstrate that while μTransfer exhibits limited scalability in large-scale settings, the Fitting paradigm—grounded in the derived scaling law—offers superior scalability and practicality, providing a systematic guideline for hyperparameter tuning in industrial-scale pretraining.

hyperparameter optimizationlarge-scale pre-traininglearning rate

This work addresses the long-standing gap in systematic research on pretraining—a phase that fundamentally determines a model’s capability ceiling—hampered by industrial opacity and academic compute constraints. Leveraging industrial-scale computational resources and full scientific autonomy, we propose Data Darwinism, a novel framework featuring an L0–L9 taxonomy of data processing depth, establishing data curation rigor as a critical dimension alongside data volume. Through two-stage adaptive curriculum learning, we train a 3B-parameter model on 8 trillion tokens and conduct over 200 ablation studies. Our findings reveal domain-specific saturation dynamics and a combinatorial balancing mechanism, demonstrating that principled data processing substantially enhances model performance while preventing catastrophic degradation. The entire pipeline is open-sourced to foster cumulative progress in the science of pretraining.

data processinglarge language modelspretraining

STEP: Staged Parameter-Efficient Pre-training for Large Language Models

Apr 05, 2025
KY
Kazuki Yano
🏛️ Tohoku University | Langsmith Inc. | RIKEN | NII LLMC

To address the prominent GPU memory bottleneck in large language model (LLM) pretraining, this paper proposes the Staged Parameter-Efficient Training (SPET) framework. SPET is the first to deeply integrate parameter-efficient fine-tuning techniques—such as LoRA—into the *entire* pretraining pipeline, synergistically combining gradient checkpointing with staged architectural expansion to enable dynamic model growth and on-demand memory optimization. Implemented in PyTorch, SPET introduces a memory-aware training scheduler that reduces peak GPU memory consumption by up to 53.9% versus full-parameter baselines, while preserving pretraining performance. Downstream task performance after instruction tuning remains unchanged. The core contribution lies in bridging the paradigmatic divide between standard pretraining and parameter-efficient adaptation, establishing a scalable, memory-efficient, and unified pretraining paradigm.

Integrates efficient tuning with model growthMaintains performance with less memoryReduces memory use in LLM pre-training

Latest Papers

What's happening recently
View more

This work addresses the challenge of sparse and noisy observational data in few-shot, large-scale decision-making problems by introducing the pretraining–fine-tuning paradigm to this setting for the first time. The authors propose a problem-specific Transformer architecture that leverages domain knowledge to generate synthetic data for pretraining, followed by fine-tuning on a small amount of real-world data. Theoretically, they establish the first non-asymptotic generalization error bound, elucidating the synergistic mechanism between pretraining and fine-tuning and revealing a scaling law for fine-tuning. Empirically, high-capacity models effectively learn structural priors from synthetic data and adapt efficiently to real environments, with decision performance improving significantly as the instance scale grows.

cross-instance learninglarge-scale optimizationnoisy observations

Midtraining Bridges Pretraining and Posttraining Distributions

Oct 16, 2025
EL
Emmy Liu
🏛️ Carnegie Mellon University

Pretraining and instruction tuning exhibit syntactic and task-distribution mismatches, leading to catastrophic forgetting of domain-specific knowledge—particularly in mathematics and code. Method: We introduce high-quality instruction data during the late pretraining phase (“midtraining”) and conduct controlled ablation studies on models trained from scratch, using diverse supervised fine-tuning datasets. Contribution/Results: We provide the first empirical evidence that midtraining functions as an effective domain adaptation technique, substantially mitigating knowledge forgetting in mathematical and programming domains. Its efficacy depends primarily on the timing of intervention—not on the proportion of instruction data mixed into pretraining. Under equal data budgets, midtraining achieves significantly lower domain-specific validation loss compared to continued pretraining. Our findings deliver causal, stage-level insights into training dynamics, establishing midtraining as a principled strategy for aligning pretraining with downstream task distributions.

It reduces syntactic disparities between pretraining and fine-tuning domainsMidtraining bridges pretraining and posttraining data distribution gapsMidtraining outperforms continued pretraining by minimizing catastrophic forgetting

This study challenges the common assumption that models exhibiting similar performance after supervised fine-tuning (SFT) are functionally equivalent, by demonstrating that the data used in the final stage of pretraining critically influences subsequent alignment behavior. Through controlled experiments—where only the last 500 million tokens of pretraining data are varied while keeping SFT and post-training procedures identical—the authors show that ending pretraining with safety-oriented text significantly preserves a model’s ability to refuse harmful requests, an effect absent with other data types. This finding is replicated across another model family, revealing for the first time that late-stage pretraining data selectively shapes how models evolve during preference optimization and reinforcement learning. The results question evaluation practices that rely solely on post-SFT performance as a proxy for alignment capability.

alignmentmodel checkpointpost-training

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

Hot Scholars

RN

Rodrigo Nogueira

Founder and CEO of Maritaca AI
Deep LearningNatural Language ProcessingInformation Retrieval
TS

Thales Sales Almeida

Student, Unicamp
Information retrievalMachine learningDeep learningGenerative models
JZ

Jingren Zhou

Alibaba Group, Microsoft
Cloud ComputingLarge Scale Distributed SystemsMachine LearningQuery Processing
RJ

Raviraj Joshi

Indian Institute of Technology Madras
computer sciencemachine learningnatural language processing
JK

Jenny Kunz

Linköping University
Natural Language Processing