fine-tune and retrain models

Designs and implements end-to-end training and retraining pipelines that pretrain, fine‑tune, adapt, or rebuild models (including pretrained transformers, regression and tabular models) and produce baseline and comparative checkpoints. Builds scalable and optimized training workflows that apply supervised learning, regularization and robustness techniques, hyperparameter tuning, validation, evaluation and benchmarking to train models at scale and measure or improve their performance.

fine-tuneandretrainmodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.42
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Update Your Transformer to the Latest Release: Re-Basin of Task Vectors

May 28, 2025
FR
Filippo Rinaldi
🏛️ AImageLab | University of Modena and Reggio Emilia | Sapienza University of Rome

When pretrained foundation models are updated, existing fine-tuned models become obsolete, necessitating efficient knowledge transfer mechanisms—especially under constraints of no access to original training data or computational resources for retraining. Method: We propose a training-free, data-free fine-tuning knowledge transfer method, the first to adapt the *re-basin* paradigm to Transformer architectures. Our approach introduces a spectral-theory-driven, two-level weight rearrangement scheme: (i) attention head permutation, (ii) intra-head parameter alignment, and (iii) task-vector rebasing. Crucially, it resolves structural inconsistencies induced by residual connections and multi-head attention. Results: The method enables zero-shot, zero-step adaptation of legacy fine-tuned models to updated pretrained backbones across vision and language tasks, fully recovering original performance without any gradient updates—eliminating the need for costly retraining.

Apply model re-basin principles to Transformer architecturesEnable data-free knowledge transfer between pretrained backbonesTransfer fine-tuning to new model releases without retraining

AdaPtis: Reducing Pipeline Bubbles with Adaptive Pipeline Parallelism on Heterogeneous Models

Sep 28, 2025
JG
Jihu Guo
🏛️ Fudan University | Shanghai AI Laboratory | Hong Kong University of Science and Technology | SenseTime | Tsinghua University | Chinese University of Hong Kong | Sensetime Research

To address severe pipeline bubbles and limited throughput in heterogeneous large-model training, this paper proposes AdaPtis, an adaptive pipeline parallelism system. Methodologically, AdaPtis introduces (1) a generalizable pipeline performance model; (2) the first joint optimization of model partitioning, device placement, and micro-batch scheduling; and (3) a unified pipeline executor supporting diverse parallelism strategies. Experiments on representative heterogeneous hardware configurations demonstrate that AdaPtis achieves an average 1.42× speedup over Megatron-LM’s I-1F1B baseline, with peak improvements reaching 2.14×. These gains translate into significantly enhanced training efficiency and improved hardware resource utilization, without compromising model accuracy or training stability.

Co-optimizes model partition, placement and workload schedulingImproves training efficiency through adaptive pipeline parallelismReduces pipeline bubbles in heterogeneous LLM training

STEP: Staged Parameter-Efficient Pre-training for Large Language Models

Apr 05, 2025
KY
Kazuki Yano
🏛️ Tohoku University | Langsmith Inc. | RIKEN | NII LLMC

To address the prominent GPU memory bottleneck in large language model (LLM) pretraining, this paper proposes the Staged Parameter-Efficient Training (SPET) framework. SPET is the first to deeply integrate parameter-efficient fine-tuning techniques—such as LoRA—into the *entire* pretraining pipeline, synergistically combining gradient checkpointing with staged architectural expansion to enable dynamic model growth and on-demand memory optimization. Implemented in PyTorch, SPET introduces a memory-aware training scheduler that reduces peak GPU memory consumption by up to 53.9% versus full-parameter baselines, while preserving pretraining performance. Downstream task performance after instruction tuning remains unchanged. The core contribution lies in bridging the paradigmatic divide between standard pretraining and parameter-efficient adaptation, establishing a scalable, memory-efficient, and unified pretraining paradigm.

Integrates efficient tuning with model growthMaintains performance with less memoryReduces memory use in LLM pre-training

Universal Checkpointing: Efficient and Flexible Checkpointing for Large Scale Distributed Training

Jun 27, 2024
XL
Xinyu Lian
🏛️ University of Illinois at Urbana-Champaign | Microsoft | StasoSphere

In large-scale DNN distributed training, checkpointing is tightly coupled with model parallelism strategies and hardware topology, severely limiting fault tolerance and elastic scalability. To address this, we propose the “distributed storage, unified loading” paradigm: during saving, model parameters are stored in a distributed representation aligned with the current parallel configuration; during restoration, they are uniformly reconstructed into a logically consistent parameter view. We design a universal checkpoint format—incorporating merged parameter representations and mapping metadata—a Universal Checkpoint Language (UCL), and an on-demand state reconstruction mechanism, achieving, for the first time, full decoupling of checkpointing from parallel configurations. Evaluated on LLaMA, Bloom, and other mainstream large models under diverse parallelism paradigms—including tensor parallelism (TP), pipeline parallelism (PP), data parallelism (DP), and context parallelism (CP)—our approach reduces post-failure recovery time by 12–28% on average, significantly enhancing cross-configuration portability and system robustness.

Decouples checkpoint structure from hardware configurationsEnables reconfigurable parallelism in large-scale DNN trainingSupports flexible mapping of checkpoint state to parallelism strategies

Latest Papers

What's happening recently
View more

This study challenges the common assumption that models exhibiting similar performance after supervised fine-tuning (SFT) are functionally equivalent, by demonstrating that the data used in the final stage of pretraining critically influences subsequent alignment behavior. Through controlled experiments—where only the last 500 million tokens of pretraining data are varied while keeping SFT and post-training procedures identical—the authors show that ending pretraining with safety-oriented text significantly preserves a model’s ability to refuse harmful requests, an effect absent with other data types. This finding is replicated across another model family, revealing for the first time that late-stage pretraining data selectively shapes how models evolve during preference optimization and reinforcement learning. The results question evaluation practices that rely solely on post-SFT performance as a proxy for alignment capability.

alignmentmodel checkpointpost-training

Hot Scholars

ZS

Zezhi Shao

Institute of Computing Technology, Chinese Academy of Sciences
Time Series ForecastingSpatial-Temporal Data MiningGraph Data Mining
GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
HZ

Hongyu Zhang

Chongqing University
Software EngineeringMining Software RepositoriesData-driven Software EngineeringSoftware Analytics
OR

Olga Russakovsky

Associate Professor, Princeton University
Computer vision
YF

Yanjie Fu

Associate Professor at School of Computing and AI, Arizona State University
Artificial IntelligenceAI4DataSpatiotemporal IntelligenceSim2Decision