Score
Designs and implements end-to-end training and retraining pipelines that pretrain, fine‑tune, adapt, or rebuild models (including pretrained transformers, regression and tabular models) and produce baseline and comparative checkpoints. Builds scalable and optimized training workflows that apply supervised learning, regularization and robustness techniques, hyperparameter tuning, validation, evaluation and benchmarking to train models at scale and measure or improve their performance.
This work addresses the fragmented landscape of post-training adaptation techniques, which suffer from inconsistent terminology and a lack of unified comparative or governance frameworks. To resolve this, the paper introduces the first six-dimensional taxonomy—spanning mechanism, objective, data requirements, persistence, structural scope, and model type—that systematically integrates mainstream approaches such as fine-tuning, retrieval augmentation, prompt engineering, model editing, and machine unlearning. This framework clarifies conceptual boundaries and reveals evolutionary and compositional relationships among methods. Beyond standardizing terminology, it enables standardized technical documentation, model change tracking, and AI governance analysis. The study further identifies critical challenges, including evaluation rigor, reproducibility, continual adaptation, multimodal alignment, and governance-aware workflows.
When pretrained foundation models are updated, existing fine-tuned models become obsolete, necessitating efficient knowledge transfer mechanisms—especially under constraints of no access to original training data or computational resources for retraining. Method: We propose a training-free, data-free fine-tuning knowledge transfer method, the first to adapt the *re-basin* paradigm to Transformer architectures. Our approach introduces a spectral-theory-driven, two-level weight rearrangement scheme: (i) attention head permutation, (ii) intra-head parameter alignment, and (iii) task-vector rebasing. Crucially, it resolves structural inconsistencies induced by residual connections and multi-head attention. Results: The method enables zero-shot, zero-step adaptation of legacy fine-tuned models to updated pretrained backbones across vision and language tasks, fully recovering original performance without any gradient updates—eliminating the need for costly retraining.
To address severe pipeline bubbles and limited throughput in heterogeneous large-model training, this paper proposes AdaPtis, an adaptive pipeline parallelism system. Methodologically, AdaPtis introduces (1) a generalizable pipeline performance model; (2) the first joint optimization of model partitioning, device placement, and micro-batch scheduling; and (3) a unified pipeline executor supporting diverse parallelism strategies. Experiments on representative heterogeneous hardware configurations demonstrate that AdaPtis achieves an average 1.42× speedup over Megatron-LM’s I-1F1B baseline, with peak improvements reaching 2.14×. These gains translate into significantly enhanced training efficiency and improved hardware resource utilization, without compromising model accuracy or training stability.
To address the prominent GPU memory bottleneck in large language model (LLM) pretraining, this paper proposes the Staged Parameter-Efficient Training (SPET) framework. SPET is the first to deeply integrate parameter-efficient fine-tuning techniques—such as LoRA—into the *entire* pretraining pipeline, synergistically combining gradient checkpointing with staged architectural expansion to enable dynamic model growth and on-demand memory optimization. Implemented in PyTorch, SPET introduces a memory-aware training scheduler that reduces peak GPU memory consumption by up to 53.9% versus full-parameter baselines, while preserving pretraining performance. Downstream task performance after instruction tuning remains unchanged. The core contribution lies in bridging the paradigmatic divide between standard pretraining and parameter-efficient adaptation, establishing a scalable, memory-efficient, and unified pretraining paradigm.
In large-scale DNN distributed training, checkpointing is tightly coupled with model parallelism strategies and hardware topology, severely limiting fault tolerance and elastic scalability. To address this, we propose the “distributed storage, unified loading” paradigm: during saving, model parameters are stored in a distributed representation aligned with the current parallel configuration; during restoration, they are uniformly reconstructed into a logically consistent parameter view. We design a universal checkpoint format—incorporating merged parameter representations and mapping metadata—a Universal Checkpoint Language (UCL), and an on-demand state reconstruction mechanism, achieving, for the first time, full decoupling of checkpointing from parallel configurations. Evaluated on LLaMA, Bloom, and other mainstream large models under diverse parallelism paradigms—including tensor parallelism (TP), pipeline parallelism (PP), data parallelism (DP), and context parallelism (CP)—our approach reduces post-failure recovery time by 12–28% on average, significantly enhancing cross-configuration portability and system robustness.
This study challenges the common assumption that models exhibiting similar performance after supervised fine-tuning (SFT) are functionally equivalent, by demonstrating that the data used in the final stage of pretraining critically influences subsequent alignment behavior. Through controlled experiments—where only the last 500 million tokens of pretraining data are varied while keeping SFT and post-training procedures identical—the authors show that ending pretraining with safety-oriented text significantly preserves a model’s ability to refuse harmful requests, an effect absent with other data types. This finding is replicated across another model family, revealing for the first time that late-stage pretraining data selectively shapes how models evolve during preference optimization and reinforcement learning. The results question evaluation practices that rely solely on post-SFT performance as a proxy for alignment capability.