pretraining

Designs, builds, and manages the initial training stage of machine learning models performed on large corpora (often unlabeled) — selecting pretraining objectives, datasets, model architectures, optimization and regularization schedules, and compute/infrastructure to produce transferable representations. Evaluates and analyzes the resulting learned representations, scaling behavior, and downstream transfer or fine‑tuning performance to guide pretraining strategy.

pretraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$213K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of sparse and noisy observational data in few-shot, large-scale decision-making problems by introducing the pretraining–fine-tuning paradigm to this setting for the first time. The authors propose a problem-specific Transformer architecture that leverages domain knowledge to generate synthetic data for pretraining, followed by fine-tuning on a small amount of real-world data. Theoretically, they establish the first non-asymptotic generalization error bound, elucidating the synergistic mechanism between pretraining and fine-tuning and revealing a scaling law for fine-tuning. Empirically, high-capacity models effectively learn structural priors from synthetic data and adapt efficiently to real environments, with decision performance improving significantly as the instance scale grows.

cross-instance learninglarge-scale optimizationnoisy observations

In large-scale pretraining, learning rate scheduling critically influences both training efficiency and model performance. This work proposes two paradigms—Fitting and Transfer. The Fitting paradigm establishes, for the first time, a scaling law for learning rate search factors, reducing hyperparameter tuning complexity from O(n³) to O(n·C_D·C_η). The Transfer paradigm extends μTransfer to Mixture-of-Experts (MoE) architectures and generalizes it across multiple hyperparameter dimensions, including depth, weight decay, and token length. Empirical results demonstrate that while μTransfer exhibits limited scalability in large-scale settings, the Fitting paradigm—grounded in the derived scaling law—offers superior scalability and practicality, providing a systematic guideline for hyperparameter tuning in industrial-scale pretraining.

hyperparameter optimizationlarge-scale pre-traininglearning rate

STEP: Staged Parameter-Efficient Pre-training for Large Language Models

Apr 05, 2025
KY
Kazuki Yano
🏛️ Tohoku University | Langsmith Inc. | RIKEN | NII LLMC

To address the prominent GPU memory bottleneck in large language model (LLM) pretraining, this paper proposes the Staged Parameter-Efficient Training (SPET) framework. SPET is the first to deeply integrate parameter-efficient fine-tuning techniques—such as LoRA—into the *entire* pretraining pipeline, synergistically combining gradient checkpointing with staged architectural expansion to enable dynamic model growth and on-demand memory optimization. Implemented in PyTorch, SPET introduces a memory-aware training scheduler that reduces peak GPU memory consumption by up to 53.9% versus full-parameter baselines, while preserving pretraining performance. Downstream task performance after instruction tuning remains unchanged. The core contribution lies in bridging the paradigmatic divide between standard pretraining and parameter-efficient adaptation, establishing a scalable, memory-efficient, and unified pretraining paradigm.

Integrates efficient tuning with model growthMaintains performance with less memoryReduces memory use in LLM pre-training

This study investigates the intrinsic synergy and trade-offs between pretraining and fine-tuning in large language models (LLMs). Methodologically, we propose a multi-stage fine-tuning analysis framework leveraging intermediate pretraining checkpoints, systematically evaluating capability improvement, adaptation to new knowledge, retention of prior knowledge, and prompt robustness across 18 diverse datasets. Key findings are: (1) continued pretraining implicitly enhances downstream fine-tuning performance; (2) fine-tuning yields substantial gains on weak-task capabilities but induces domain-specific knowledge forgetting; (3) fine-tuning exacerbates prompt sensitivity, whereas additional pretraining effectively mitigates this effect. Crucially, we quantitatively demonstrate the reversibility of both knowledge forgetting and prompt sensitivity—establishing that pretraining quality fundamentally bounds fine-tuning efficacy. Our work provides the first reproducible empirical guidelines and standardized evaluation protocols for optimizing the pretraining–fine-tuning pipeline.

Fine-tuningLarge Language ModelsTransfer Learning

Deep Learning and Machine Learning, Advancing Big Data Analytics and Management: Tensorflow Pretrained Models

Sep 20, 2024
KC
Keyu Chen
🏛️ Georgia Institute of Technology | Indiana University | Kyoto University | AppCubic | Rutgers University | Purdue University | University of Wisconsin-Madison | National Taiwan Normal University

High barriers to adopting pre-trained models and a lack of empirical guidance for strategy selection hinder practical deployment in few-shot image classification and object detection. Method: We systematically compare linear probing versus fine-tuning across ResNet, MobileNet, and EfficientNet, and propose an end-to-end TensorFlow framework integrating multi-scale feature-space visualization (PCA, t-SNE, UMAP) to unify analysis of representation evolution. Contribution/Results: Linear probing significantly outperforms fine-tuning under extreme data scarcity (≤100 samples per class) while accelerating training by 3–5×. The framework enables high-accuracy, rapid deployment (<1 hour for fine-tuning) on standard benchmarks (ImageNet-1K, CIFAR-100), balancing beginner-friendly usability with expert-level extensibility. It bridges the gap between theoretical representation analysis and real-world engineering practice.

Comparing linear probing versus fine-tuning approaches in transfer learningExploring TensorFlow pre-trained models for image classification tasksProviding practical guidance and code examples for deep learning implementation

Latest Papers

What's happening recently
View more

This work addresses the long-standing gap in systematic research on pretraining—a phase that fundamentally determines a model’s capability ceiling—hampered by industrial opacity and academic compute constraints. Leveraging industrial-scale computational resources and full scientific autonomy, we propose Data Darwinism, a novel framework featuring an L0–L9 taxonomy of data processing depth, establishing data curation rigor as a critical dimension alongside data volume. Through two-stage adaptive curriculum learning, we train a 3B-parameter model on 8 trillion tokens and conduct over 200 ablation studies. Our findings reveal domain-specific saturation dynamics and a combinatorial balancing mechanism, demonstrating that principled data processing substantially enhances model performance while preventing catastrophic degradation. The entire pipeline is open-sourced to foster cumulative progress in the science of pretraining.

data processinglarge language modelspretraining

This study addresses the challenge of extracting machine learning pipeline stages, which is constrained by domain diversity and where existing methods rely on manual annotation or limited classifiers. This work systematically investigates, for the first time, the potential of small language models (SLMs) to parse ML pipeline structures leveraging their inherent code comprehension capabilities without fine-tuning, employing Cochran’s Q test, McNemar’s test, and goodness-of-fit evaluations for rigorous assessment. The findings indicate that while SLMs demonstrate robust performance, they do not surpass existing classifiers; however, the core contribution lies in revealing that different classification approaches significantly influence practical insights. Despite the limitation of high inference costs, this research establishes a novel paradigm for automated ML structure parsing.

Code ClassificationMachine Learning PipelinesReverse Engineering

This work addresses the critical challenge in continual learning of mitigating catastrophic forgetting during downstream fine-tuning while preserving capabilities acquired during upstream training. The authors propose treating “robustness to subsequent fine-tuning” as a first-class objective in upstream training and systematically investigate data scheduling strategies across a three-stage pipeline—pretraining, post-training, and downstream fine-tuning. Their key finding is that early exposure to post-training data during pretraining—termed “early data exposure”—consistently outperforms pure post-training or conventional mixing strategies, yielding superior trade-offs between upstream knowledge retention and downstream task performance across model scales from 135M to 1B parameters. This approach complements regularization techniques such as replay and Dropout and, under fixed compute budgets, reveals an optimal data allocation scheme.

catastrophic forgettingearly exposurefine-tuning

This work addresses the challenge of selecting effective fine-tuning strategies for encoder-decoder pre-trained language models in generation and question-answering tasks. It proposes the Match Task to Objective (MTO) framework, which establishes the first systematic alignment mechanism between downstream tasks and pre-training objectives. MTO automatically constructs training data and prompt templates that are consistent with the original pre-training objective and extends this alignment to soft prompt tuning, thereby enabling precise task–objective matching. Experimental results demonstrate that MTO achieves over 120% performance improvement under few-shot settings compared to existing methods, significantly outperforms strong baselines in full-data scenarios, and substantially enhances the effectiveness of prompt tuning.

commonsense reasoningencoder-decoder modelspre-training objectives

Hot Scholars

HZ

Hao Zhao

Tsinghua University
Computer Vision
FX

Fangzhi Xu

Xi'an Jiaotong University | Nanyang Technological University
Large Language ModelsSelf-TrainingReasoningGUI Agents
YZ

Yifan Zhu

Beijing University of Posts and Telecommunications
PEFT of LLMsGraph RAGGraph mining
CR

Cameron R. Jones

Postdoc, UC San Diego
large language modelsturing testsocial intelligence
QL

Qika Lin

National University of Singapore | NTU | XJTU | BIT
Knowledge ReasoningNeurosymbolic AIMulti-modalRobustness & Security