staged multi-task training

Design, implement, and evaluate multi-stage training pipelines that partition learning into ordered stages with distinct objectives, data sampling strategies, and curricula so capabilities are acquired and refined progressively; this includes capability‑oriented objectives, adaptive data‑scheduling, stagewise distillation and fine‑tuning, and monitoring student performance per stage. It also covers methods for coordinating specialized components across stages—e.g., two‑stage pretrain/finetune procedures, frozen‑expert weight learning, mixture‑of‑experts training, and domain‑adaptive fine‑tuning—to shift model behavior or reduce synthetic‑to‑real and dataset‑mismatch issues.

stagedmulti-tasktraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Mid-Training of Large Language Models: A Survey

Oct 08, 2025
KM
Kaixiang Mo
🏛️ Shopee

This work addresses the lack of a unified theoretical framework and empirical analysis for the mid-training phase—intervening between pretraining and downstream fine-tuning—in large language models (LLMs). We propose the first systematic taxonomy covering data distribution evolution, learning rate annealing scheduling, and long-context extension; explain mid-training efficacy through gradient noise suppression, information bottleneck alleviation, and curriculum learning; and establish a standardized evaluation benchmark with reproducible training guidelines. Experiments demonstrate that mid-training significantly improves model generalization and subsequent fine-tuning efficiency. Our findings provide a structured methodological foundation for continuous LLM capability evolution and highlight open challenges—including data-optimization co-design and dynamic context adaptation—that warrant further investigation.

Addressing diminishing returns and convergence instability in LLM trainingEstablishing unified taxonomy and benchmarks for mid-training evaluationInvestigating intermediate training stage between pre-training and fine-tuning

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the unclear trade-offs among diverse capabilities—such as general visual understanding, structured reasoning, and fine-grained OCR—in multimodal instruction tuning under mixed data regimes, particularly the lack of systematic investigation into how data organization influences these trade-offs. Treating data scheduling as a first-order design variable while holding model architecture and optimization settings fixed, the work compares four strategies: direct mixing, curriculum learning, balanced sampling, and reverse curriculum. Results demonstrate that curriculum-based training—sequencing tasks from general comprehension to specialized skills—achieves superior overall performance and structured reasoning while accelerating convergence. Balanced sampling improves OCR accuracy at the cost of capability imbalance, whereas reverse curriculum degrades performance and induces optimization instability. The findings highlight the critical role of training sequence in shaping the capability distribution of multimodal models.

capability trade-offsdata organizationmultimodal instruction tuning

This work addresses the long-standing gap in systematic research on pretraining—a phase that fundamentally determines a model’s capability ceiling—hampered by industrial opacity and academic compute constraints. Leveraging industrial-scale computational resources and full scientific autonomy, we propose Data Darwinism, a novel framework featuring an L0–L9 taxonomy of data processing depth, establishing data curation rigor as a critical dimension alongside data volume. Through two-stage adaptive curriculum learning, we train a 3B-parameter model on 8 trillion tokens and conduct over 200 ablation studies. Our findings reveal domain-specific saturation dynamics and a combinatorial balancing mechanism, demonstrating that principled data processing substantially enhances model performance while preventing catastrophic degradation. The entire pipeline is open-sourced to foster cumulative progress in the science of pretraining.

data processinglarge language modelspretraining

STEP: Staged Parameter-Efficient Pre-training for Large Language Models

Apr 05, 2025
KY
Kazuki Yano
🏛️ Tohoku University | Langsmith Inc. | RIKEN | NII LLMC

To address the prominent GPU memory bottleneck in large language model (LLM) pretraining, this paper proposes the Staged Parameter-Efficient Training (SPET) framework. SPET is the first to deeply integrate parameter-efficient fine-tuning techniques—such as LoRA—into the *entire* pretraining pipeline, synergistically combining gradient checkpointing with staged architectural expansion to enable dynamic model growth and on-demand memory optimization. Implemented in PyTorch, SPET introduces a memory-aware training scheduler that reduces peak GPU memory consumption by up to 53.9% versus full-parameter baselines, while preserving pretraining performance. Downstream task performance after instruction tuning remains unchanged. The core contribution lies in bridging the paradigmatic divide between standard pretraining and parameter-efficient adaptation, establishing a scalable, memory-efficient, and unified pretraining paradigm.

Integrates efficient tuning with model growthMaintains performance with less memoryReduces memory use in LLM pre-training

Mitigating Forgetting in LLM Supervised Fine-Tuning and Preference Learning

Oct 20, 2024
HF
Heshan Fernando
🏛️ Rensselaer Polytechnic Institute | IBM Research

Sequential supervised fine-tuning (SFT) followed by preference learning (e.g., DPO or RLHF) in large language model post-training induces catastrophic forgetting, degrading SFT task performance while optimizing for preference alignment. Method: This work theoretically establishes the suboptimality of sequential training and proposes the first jointly optimized framework with provable convergence guarantees. It unifies SFT and DPO objectives via a scalable multi-objective loss function and performs joint gradient updates—without additional computational overhead. Contribution/Results: The method enables effective knowledge fusion across both stages. Experiments demonstrate a 23% improvement in SFT task retention and a 9.7% gain in preference alignment accuracy over sequential baselines, while maintaining comparable computational cost.

Mitigate forgetting in LLM post-trainingOptimize SFT and RLHF/DPO trade-offPropose joint post-training framework

Latest Papers

What's happening recently
View more

This work addresses the fragmented landscape of post-training adaptation techniques, which suffer from inconsistent terminology and a lack of unified comparative or governance frameworks. To resolve this, the paper introduces the first six-dimensional taxonomy—spanning mechanism, objective, data requirements, persistence, structural scope, and model type—that systematically integrates mainstream approaches such as fine-tuning, retrieval augmentation, prompt engineering, model editing, and machine unlearning. This framework clarifies conceptual boundaries and reveals evolutionary and compositional relationships among methods. Beyond standardizing terminology, it enables standardized technical documentation, model change tracking, and AI governance analysis. The study further identifies critical challenges, including evaluation rigor, reproducibility, continual adaptation, multimodal alignment, and governance-aware workflows.

AI governancefoundation modelsmodel modification

This work addresses the critical challenge in continual learning of mitigating catastrophic forgetting during downstream fine-tuning while preserving capabilities acquired during upstream training. The authors propose treating “robustness to subsequent fine-tuning” as a first-class objective in upstream training and systematically investigate data scheduling strategies across a three-stage pipeline—pretraining, post-training, and downstream fine-tuning. Their key finding is that early exposure to post-training data during pretraining—termed “early data exposure”—consistently outperforms pure post-training or conventional mixing strategies, yielding superior trade-offs between upstream knowledge retention and downstream task performance across model scales from 135M to 1B parameters. This approach complements regularization techniques such as replay and Dropout and, under fixed compute budgets, reveals an optimal data allocation scheme.

catastrophic forgettingearly exposurefine-tuning

Existing large-scale model training systems struggle to flexibly compose diverse parallelization strategies, often relying on manual expert tuning and lacking generality. This work proposes a programmable distributed training system that enables users to declaratively specify composite parallelism strategies—such as data, pipeline, and expert parallelism—through model annotations and scheduling directives. These specifications are compiled via a unified intermediate representation (IR) into device-level execution plans, fully decoupling strategy definition from runtime execution over a global compute-communication DAG. The system is the first to support automatic compilation of user-defined composite strategies, matching the performance of established approaches like ZeRO while significantly improving both performance and memory efficiency in complex scenarios such as DeepSeek-V3’s DualPipe.

distributed trainingflexibilitymodel parallelism

Hot Scholars

XW

Xinggang Wang

Professor, Huazhong University of Science and Technology
Artificial IntelligenceComputer VisionAutonomous DrivingObject Detection
JQ

Jie Qin

Professor, Nanjing University of Aeronautics and Astronautics
Computer VisionMachine LearningPattern Recognition
XS

Xing Sun

Tencent Youtu Lab
LLMMLLMAgent
PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
YZ

Yuanxing Zhang

Kuaishou Technology
Recommender SystemLarge Language ModelVideo Understanding