train models from scratch

Designs and executes complete model training and retraining pipelines that begin from random weight initialization (training from scratch), including architecture setup, data preprocessing, optimizer and hyperparameter configuration, checkpointing, and evaluation. Builds workflows to retrain models from a base state (retraining from scratch) to compare against pretrained variants, implement approximate unlearning strategies, and analyze training dynamics, performance, and retention.

trainmodelsfromscratch

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Universal Checkpointing: Efficient and Flexible Checkpointing for Large Scale Distributed Training

Jun 27, 2024
XL
Xinyu Lian
🏛️ University of Illinois at Urbana-Champaign | Microsoft | StasoSphere

In large-scale DNN distributed training, checkpointing is tightly coupled with model parallelism strategies and hardware topology, severely limiting fault tolerance and elastic scalability. To address this, we propose the “distributed storage, unified loading” paradigm: during saving, model parameters are stored in a distributed representation aligned with the current parallel configuration; during restoration, they are uniformly reconstructed into a logically consistent parameter view. We design a universal checkpoint format—incorporating merged parameter representations and mapping metadata—a Universal Checkpoint Language (UCL), and an on-demand state reconstruction mechanism, achieving, for the first time, full decoupling of checkpointing from parallel configurations. Evaluated on LLaMA, Bloom, and other mainstream large models under diverse parallelism paradigms—including tensor parallelism (TP), pipeline parallelism (PP), data parallelism (DP), and context parallelism (CP)—our approach reduces post-failure recovery time by 12–28% on average, significantly enhancing cross-configuration portability and system robustness.

Decouples checkpoint structure from hardware configurationsEnables reconfigurable parallelism in large-scale DNN trainingSupports flexible mapping of checkpoint state to parallelism strategies

Pre-trained models struggle to adapt efficiently to diverse deployment sizes (e.g., varying depth or width), necessitating size-specific initialization strategies. Method: This paper proposes a multi-task-inspired framework for variable-size model initialization. Its core innovation is a size-agnostic shared weight template—constructed via Learngene-based knowledge distillation with Kronecker-structured constraints—and lightweight learnable scalers that enable consistent cross-size initialization. The template-scaler co-design supports zero-shot transfer to diverse downstream tasks, while scalers require only minimal data for adaptation. Contribution/Results: Experiments demonstrate state-of-the-art initialization performance across multiple depth- and width-varied architectures. The method significantly enhances few-shot adaptability and cross-task generalization, offering a scalable, data-efficient solution for deploying models of heterogeneous sizes without retraining from scratch.

Addresses initialization of variable-sized modelsEnsures consistent initialization across different model sizesOvercomes limitations of pre-trained model sizes

PISTOL: Dataset Compilation Pipeline for Structural Unlearning of LLMs

Jun 24, 2024
XQ
Xinchi Qiu
🏛️ University of Cambridge | UCL | Meta

Existing unlearning methods for large language models treat forget samples as isolated instances, neglecting their structural interdependencies—such as logical entailments in knowledge graphs, batch-level correlations, and domain-level shifts—leading to incomplete or cascading forgetting. Method: We propose *structural unlearning*, a novel paradigm grounded in knowledge graph semantics. We introduce StructUnlearn, the first knowledge graph–driven structural unlearning benchmark, featuring multi-hop fact perturbation, domain shift simulation, and batch dependency modeling, supported by a reproducible synthetic data generation pipeline. We establish a cross-model evaluation framework using Llama2-7B and Mistral-7B. Contributions/Results: Our systematic evaluation exposes critical failure modes of four mainstream unlearning approaches under interconnected fact scenarios. We provide the first empirical evidence that pretraining model selection significantly impacts forgetting robustness. Moreover, we rigorously characterize the fundamental trade-off between utility preservation and complete forgetting—highlighting structural dependencies as a key determinant.

Analyzes unlearning difficulty with knowledge graph density and domain skew.Explores impact of data inter-connectivity on LLM unlearning.Proposes PISTOL for structural dataset compilation and evaluation.

Towards Scalable Exact Machine Unlearning Using Parameter-Efficient Fine-Tuning

Jun 24, 2024
SB
Somnath Basu Roy Chowdhury
🏛️ UNC Chapel Hill | Google DeepMind | Columbia University | Independent | Google Research

This paper addresses the challenges of “exact removal of specific training samples” in machine unlearning—namely, high computational overhead from retraining, significant latency, and degradation in model performance. To this end, we propose the Sequence-Aware Sharded and Stratified Training (S3T) framework. S3T employs hierarchical sequential training, disjoint partitioning of data subsets, and layer-wise parameter isolation, enabling theoretically rigorous, zero-loss exact unlearning via deactivation of only affected layers. It is the first method to support high-concurrency deletion requests while guaranteeing zero service interruption. Integrated with parameter-efficient fine-tuning (PEFT) and multi-sequence joint optimization, S3T achieves substantial improvements across multiple benchmarks: 92% reduction in deletion latency, <0.3% accuracy loss, 100% service availability, and formal theoretical guarantees on deletion equivalence and performance consistency.

Efficiently remove data influence without full retraining.Enhance deletion capabilities while maintaining model performance.Minimize system downtime during model component retraining.

STEP: Staged Parameter-Efficient Pre-training for Large Language Models

Apr 05, 2025
KY
Kazuki Yano
🏛️ Tohoku University | Langsmith Inc. | RIKEN | NII LLMC

To address the prominent GPU memory bottleneck in large language model (LLM) pretraining, this paper proposes the Staged Parameter-Efficient Training (SPET) framework. SPET is the first to deeply integrate parameter-efficient fine-tuning techniques—such as LoRA—into the *entire* pretraining pipeline, synergistically combining gradient checkpointing with staged architectural expansion to enable dynamic model growth and on-demand memory optimization. Implemented in PyTorch, SPET introduces a memory-aware training scheduler that reduces peak GPU memory consumption by up to 53.9% versus full-parameter baselines, while preserving pretraining performance. Downstream task performance after instruction tuning remains unchanged. The core contribution lies in bridging the paradigmatic divide between standard pretraining and parameter-efficient adaptation, establishing a scalable, memory-efficient, and unified pretraining paradigm.

Integrates efficient tuning with model growthMaintains performance with less memoryReduces memory use in LLM pre-training

Latest Papers

What's happening recently
View more

This work addresses the longstanding conflation in machine unlearning research between “untraining” and “unlearning,” which has led to ambiguous problem formulations and inadequate evaluation criteria. We formally distinguish these concepts for the first time: untraining aims to remove the influence of specific training samples, whereas true unlearning requires erasing the model’s knowledge of the entire underlying data distribution or concept those samples represent. Through theoretical formalization and a systematic review of existing literature, we establish a clear conceptual framework, reclassify current methods accordingly, and uncover critical challenges that have been overlooked. By clarifying foundational definitions, this study lays the groundwork for rigorous algorithmic evaluation, promotes standardization in the field, and delineates promising directions for future research.

concept removaldata deletionmachine unlearning

This work addresses the lack of transparent, scalable, and deeply PyTorch-integrated open-source tools for post-training large language models, which hinders research iteration and deployment efficiency. We propose a native PyTorch-based, modular post-training framework centered on the principle of “hackability,” offering composable model builders, training recipes, and a distributed training stack that support diverse fine-tuning strategies and hardware configurations. While maintaining high performance and memory efficiency, the framework significantly enhances code transparency and research flexibility. Empirical evaluations demonstrate that it matches or even surpasses mainstream tools such as Axolotl and Unsloth across multiple post-training scenarios, thereby facilitating efficient and reproducible scientific exploration.

extensibilityfine-tuninglarge language models

Hot Scholars

DL

Dahua Lin

The Chinese University of Hong Kong
computer visionmachine learningprobabilistic inferencebayesian nonparametrics
RJ

Raviraj Joshi

Indian Institute of Technology Madras
computer sciencemachine learningnatural language processing
SK

Subbarao Kambhampati

Arizona State University
Artificial IntelligenceAutomated planningLLM ReasoningHuman-AI Interaction
JE

Joshua Engels

Google Deepmind
Mechanistic InterpretabilityAI Safety
MQ

Minghui Qiu

Alibaba Group
Deep LearningTransfer LearningChatbotsNLP