state checkpointing and recovery

Design and implement mechanisms to save, restore, and manage program or model state for fault-tolerant restart and training resumption — including automated and transparent checkpoint insertion, incremental/process-delta dumps, model, activation, and gradient checkpointing, and the checkpoint/restart routines themselves. Build checkpoint-selection, monitoring, and management tools and integrate checkpoint/restart logic with runtime error handling (e.g., MPI) to realize resilient, fault-tolerant checkpointing and recovery strategies.

statecheckpointingandrecovery

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$222K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Universal Checkpointing: Efficient and Flexible Checkpointing for Large Scale Distributed Training

Jun 27, 2024
XL
Xinyu Lian
🏛️ University of Illinois at Urbana-Champaign | Microsoft | StasoSphere

In large-scale DNN distributed training, checkpointing is tightly coupled with model parallelism strategies and hardware topology, severely limiting fault tolerance and elastic scalability. To address this, we propose the “distributed storage, unified loading” paradigm: during saving, model parameters are stored in a distributed representation aligned with the current parallel configuration; during restoration, they are uniformly reconstructed into a logically consistent parameter view. We design a universal checkpoint format—incorporating merged parameter representations and mapping metadata—a Universal Checkpoint Language (UCL), and an on-demand state reconstruction mechanism, achieving, for the first time, full decoupling of checkpointing from parallel configurations. Evaluated on LLaMA, Bloom, and other mainstream large models under diverse parallelism paradigms—including tensor parallelism (TP), pipeline parallelism (PP), data parallelism (DP), and context parallelism (CP)—our approach reduces post-failure recovery time by 12–28% on average, significantly enhancing cross-configuration portability and system robustness.

Decouples checkpoint structure from hardware configurationsEnables reconfigurable parallelism in large-scale DNN trainingSupports flexible mapping of checkpoint state to parallelism strategies

Checkpoint-based rollback recovery in session programming

Dec 03, 2023
CA
C. A. Mezzina
🏛️ Università di Urbino | Università degli Studi di Firenze | University of Oxford

To address communication-intensive session programming, this paper proposes a novel session type system supporting explicit commit, rollback, and abort operations. To prevent illegal cross-participant state restoration, it introduces—within the session types framework—the first statically decidable rollback compliance check, ensuring rollbacks affect only locally accessed states and never violate inter-participant boundaries. Building on session type theory, the authors extend the session language and integrate MAUDE for design-time, type-level verification. They formally prove that the system satisfies error-freedom and progress properties. The core contributions are: (i) a safe and controllable rollback semantics grounded in session types; (ii) static enforcement of cross-participant isolation during rollback; and (iii) verifiable session recovery behavior, enabling rigorous reasoning about fault-tolerant distributed protocols.

Concurrency ControlException HandlingState Recovery

ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development

Jul 29, 2024
BW
Borui Wan
🏛️ The University of Hong Kong | ByteDance

Training large foundation models (LFMs) faces significant challenges in checkpoint management, including poor cross-framework compatibility, tight coupling with parallelization strategies, heterogeneous storage backends, and severe I/O bottlenecks. To address these, this work proposes an industrial-grade unified archival system. Its core contributions are: (1) a novel parallelism-agnostic checkpoint serialization format; (2) a full-stack I/O optimization framework integrating a dynamic resharding engine, multi-framework abstraction interfaces (PyTorch/Megatron/DeepSpeed), asynchronous high-throughput storage adapters, and a distributed I/O monitoring toolchain; and (3) runtime support for cross-parallelism resharding, multi-backend adaptivity, and rapid failure recovery. Experiments demonstrate an average 54.20× reduction in checkpoint blocking time, with peak checkpoint save and load speedups of 9.96× and 8.80×, respectively. The system has been stably deployed in production environments scaling to over one thousand GPUs.

Efficient checkpoint management for Large Foundation Models.Reduction of runtime checkpoint stalls and improved I/O efficiency.Support for multiple training frameworks and storage backends.

Manually implementing checkpoint/restart mechanisms for MPI-based scientific applications is time-consuming and heavily reliant on expert knowledge. This work proposes an automated framework that leverages the Claude Opus 4.7 large language model in conjunction with OpenCode CLI to iteratively generate code, achieving end-to-end fault tolerance injection without human intervention for the first time. Evaluated on six real-world MPI applications, the approach automatically produces valid checkpointing code within an average of 50 minutes, incurs negligible runtime overhead, and achieves recovery performance comparable to handcrafted implementations. This significantly lowers the technical barrier to resilience engineering in high-performance computing applications.

checkpointingfault toleranceMPI

Current quantum high-performance computing systems lack effective fault-tolerance and recovery mechanisms, and conventional checkpointing approaches based on quantum state preservation are fundamentally hindered by the no-cloning theorem. This work proposes a novel algorithm-level fault-tolerance framework that reframes checkpointing and recovery as problems of control flow and algorithmic state management, thereby avoiding direct storage of quantum states. Instead, it leverages mid-circuit measurements, classical feedforward, and conditional operations in dynamic quantum circuits to efficiently capture and restore program execution states. The approach enables reliable interruption and resumption for representative iterative or staged quantum algorithms—including variational eigensolvers, the Quantum Approximate Optimization Algorithm (QAOA), and time-stepping simulations—significantly enhancing the reliability and restartability of quantum computational tasks.

checkpointingfault tolerancequantum computing

Latest Papers

What's happening recently
View more

This study addresses the high implementation barrier and reliance on domain expertise associated with checkpoint/restart mechanisms in HPC scientific applications by proposing an automated fault-tolerance framework driven by large language model (LLM) coding agents. The framework establishes a generate-verify-revise closed-loop pipeline that autonomously injects fault-tolerance capabilities into MPI applications without human intervention. This work provides the first demonstration that LLM agents can efficiently perform resilience engineering given state visibility. Experimental results show that the proposed approach successfully produces 41 functional implementations, each requiring under one hour on average, while introducing negligible runtime overhead. Furthermore, the achieved fault recovery efficiency is comparable to that of manually written code, substantially reducing development costs.

checkpoint/restarthigh-performance computingLLM agents

This study addresses the miscalculation of recovery costs in checkpointing systems for coding agents, which stems from the inability to distinguish valid safety boundaries. We propose a counterfactual checkpoint advantage framework that drives dual branches toward identical logical failures and matches conditional recoveries to quantify the performance differential between saving and skipping states. This approach reveals that recovery fundamentally entails re-derivation rather than simple replay, demonstrating that traditional evaluation metrics based on preserved workload introduce severe bias. Controlled experiments on the SWE-bench Verified benchmark show that this strategy saves an average of 49.4 seconds across twelve tasks. Furthermore, the observed first-checkpoint transition rate is 0.64, whereas subsequent checkpoints yield negative values, thereby establishing a performance baseline for optimal checkpoint placement.

checkpoint advantagecheckpoint placementcoding agents

This work addresses the inefficiency of existing fault-tolerance mechanisms for large language model (LLM) agents, which rely on full restarts or application-level checkpointing and struggle to recover critical GPU state—such as KV caches and scheduler metadata—efficiently. The authors propose a device-resident, persistent kernel implementation that enables fine-grained, low-overhead checkpointing and recovery directly within the GPU’s native execution context, thereby avoiding CPU bottlenecks. By injecting incremental checkpoint logic into the PTX/SASS layer via JIT compilation, and combining GPU module load interception, lock-free ring buffers, and CXL/host-memory logging, the approach achieves framework-agnostic, transparent fault tolerance. It supports efficient dirty-page detection and incremental snapshots of key structures like KV caches and adapter pages, substantially reducing work loss and recovery latency upon failure.

checkpointingfault toleranceGPU-resident state

Hot Scholars

JF

Junfeng Fang

National University of Singapore
Model EditingAI SafetyLLM ExplainabilityAI4Science
PP

Patrick P. C. Lee

The Chinese University of Hong Kong
storage systemsnetworksdistributed systemsdependability
BN

Bogdan Nicolae

Argonne National Laboratory
High Performance ComputingAIParallel and Distributed SystemsStorage
YC

Yejin Choi

Stanford University / NVIDIA
Natural Language ProcessingDeep LearningArtificial IntelligenceCommonsense Reasoning
WX

Wayne Xin Zhao

Professor, Renmin University of China
Recommender SystemNatural Language ProcessingLarge Language Model