Score
Designs and evaluates pretraining regimes for machine learning models, specifying pretraining objectives, optimization and data-selection methods, curricula and schedules for continued/continual/continuous pretraining, and techniques for efficient pretraining. Also develops procedures and trade-offs for combining pretraining with subsequent fine-tuning to meet target performance, compute, or data-efficiency constraints.
The “mid-training” phase—occurring between pretraining and post-training—has been largely overlooked in large language model (LLM) development, despite its critical role in balancing targeted capability enhancement with preservation of foundational language modeling performance. Method: We formally define and categorize mid-training for the first time, proposing a multi-stage optimization framework encompassing data curation, curriculum learning, continued pretraining, instruction tuning, and architectural expansion. Contribution/Results: Empirical evaluation demonstrates that our approach systematically improves target capabilities—including mathematical reasoning, code generation, complex reasoning, and long-context understanding—while robustly maintaining general language modeling performance. This work establishes the first theoretical framework and practical guidelines for mid-training, enabling reproducible, controllable, and efficient LLM capability evolution and domain-specific customization.
This paper identifies a systemic issue in machine learning: preprocessing hyperparameters—such as missing-value imputation strategies—are frequently overlooked yet substantially bias model evaluation. Current practice often involves informal, post-hoc tuning of preprocessing steps, leading to optimistic performance estimates and irreproducible results. To address this, the authors formally distinguish and empirically analyze the coupling effects between algorithmic and preprocessing hyperparameters. Using a modular supervised learning workflow model, controlled variable experiments, replication of canonical case studies, and bias diagnostics, they quantify the resulting optimistic bias. Key contributions include: (1) establishing preprocessing hyperparameters as equally critical as algorithmic ones; (2) proposing formal modeling principles to eliminate informal preprocessing tuning; and (3) delivering actionable reporting guidelines for ML practitioners, thereby significantly enhancing model credibility and reproducibility.
This paper addresses the fundamental challenge of balancing catastrophic forgetting and parameter efficiency when large pre-trained models continuously adapt to dynamic task streams. To this end, we propose the first unified theoretical framework for Parameter-Efficient Continual Fine-Tuning (PECFT). Our framework systematically organizes existing approaches along three dimensions: method taxonomy, evaluation metrics, and core challenges—integrating Parameter-Efficient Fine-Tuning (PEFT) techniques (e.g., adapters, LoRA, prompt tuning) with continual learning strategies (e.g., replay, regularization, architecture expansion). Through a comprehensive review of over 100 studies, we identify key trade-offs between performance and efficiency, and pinpoint scalable memory mechanisms and task-aware parameter updates as critical research frontiers. This work bridges a significant gap at the intersection of continual learning and PEFT, providing both theoretical foundations and practical guidelines for efficient, sustainable adaptation of large language models.
In large-scale pretraining, learning rate scheduling critically influences both training efficiency and model performance. This work proposes two paradigms—Fitting and Transfer. The Fitting paradigm establishes, for the first time, a scaling law for learning rate search factors, reducing hyperparameter tuning complexity from O(n³) to O(n·C_D·C_η). The Transfer paradigm extends μTransfer to Mixture-of-Experts (MoE) architectures and generalizes it across multiple hyperparameter dimensions, including depth, weight decay, and token length. Empirical results demonstrate that while μTransfer exhibits limited scalability in large-scale settings, the Fitting paradigm—grounded in the derived scaling law—offers superior scalability and practicality, providing a systematic guideline for hyperparameter tuning in industrial-scale pretraining.
This work addresses the long-standing gap in systematic research on pretraining—a phase that fundamentally determines a model’s capability ceiling—hampered by industrial opacity and academic compute constraints. Leveraging industrial-scale computational resources and full scientific autonomy, we propose Data Darwinism, a novel framework featuring an L0–L9 taxonomy of data processing depth, establishing data curation rigor as a critical dimension alongside data volume. Through two-stage adaptive curriculum learning, we train a 3B-parameter model on 8 trillion tokens and conduct over 200 ablation studies. Our findings reveal domain-specific saturation dynamics and a combinatorial balancing mechanism, demonstrating that principled data processing substantially enhances model performance while preventing catastrophic degradation. The entire pipeline is open-sourced to foster cumulative progress in the science of pretraining.
To address the prominent GPU memory bottleneck in large language model (LLM) pretraining, this paper proposes the Staged Parameter-Efficient Training (SPET) framework. SPET is the first to deeply integrate parameter-efficient fine-tuning techniques—such as LoRA—into the *entire* pretraining pipeline, synergistically combining gradient checkpointing with staged architectural expansion to enable dynamic model growth and on-demand memory optimization. Implemented in PyTorch, SPET introduces a memory-aware training scheduler that reduces peak GPU memory consumption by up to 53.9% versus full-parameter baselines, while preserving pretraining performance. Downstream task performance after instruction tuning remains unchanged. The core contribution lies in bridging the paradigmatic divide between standard pretraining and parameter-efficient adaptation, establishing a scalable, memory-efficient, and unified pretraining paradigm.
This work addresses the challenge of sparse and noisy observational data in few-shot, large-scale decision-making problems by introducing the pretraining–fine-tuning paradigm to this setting for the first time. The authors propose a problem-specific Transformer architecture that leverages domain knowledge to generate synthetic data for pretraining, followed by fine-tuning on a small amount of real-world data. Theoretically, they establish the first non-asymptotic generalization error bound, elucidating the synergistic mechanism between pretraining and fine-tuning and revealing a scaling law for fine-tuning. Empirically, high-capacity models effectively learn structural priors from synthetic data and adapt efficiently to real environments, with decision performance improving significantly as the instance scale grows.
Pretraining and instruction tuning exhibit syntactic and task-distribution mismatches, leading to catastrophic forgetting of domain-specific knowledge—particularly in mathematics and code. Method: We introduce high-quality instruction data during the late pretraining phase (“midtraining”) and conduct controlled ablation studies on models trained from scratch, using diverse supervised fine-tuning datasets. Contribution/Results: We provide the first empirical evidence that midtraining functions as an effective domain adaptation technique, substantially mitigating knowledge forgetting in mathematical and programming domains. Its efficacy depends primarily on the timing of intervention—not on the proportion of instruction data mixed into pretraining. Under equal data budgets, midtraining achieves significantly lower domain-specific validation loss compared to continued pretraining. Our findings deliver causal, stage-level insights into training dynamics, establishing midtraining as a principled strategy for aligning pretraining with downstream task distributions.
This study challenges the common assumption that models exhibiting similar performance after supervised fine-tuning (SFT) are functionally equivalent, by demonstrating that the data used in the final stage of pretraining critically influences subsequent alignment behavior. Through controlled experiments—where only the last 500 million tokens of pretraining data are varied while keeping SFT and post-training procedures identical—the authors show that ending pretraining with safety-oriented text significantly preserves a model’s ability to refuse harmful requests, an effect absent with other data types. This finding is replicated across another model family, revealing for the first time that late-stage pretraining data selectively shapes how models evolve during preference optimization and reinforcement learning. The results question evaluation practices that rely solely on post-SFT performance as a proxy for alignment capability.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
本文研究了在固定训练预算下,如何最优分配预训练和微调的计算资源问题,通过正则化最小二乘法和梯度下降方法来解决。