Score
Design and implement mechanisms to save, restore, and manage program or model state for fault-tolerant restart and training resumption — including automated and transparent checkpoint insertion, incremental/process-delta dumps, model, activation, and gradient checkpointing, and the checkpoint/restart routines themselves. Build checkpoint-selection, monitoring, and management tools and integrate checkpoint/restart logic with runtime error handling (e.g., MPI) to realize resilient, fault-tolerant checkpointing and recovery strategies.
In large-scale DNN distributed training, checkpointing is tightly coupled with model parallelism strategies and hardware topology, severely limiting fault tolerance and elastic scalability. To address this, we propose the “distributed storage, unified loading” paradigm: during saving, model parameters are stored in a distributed representation aligned with the current parallel configuration; during restoration, they are uniformly reconstructed into a logically consistent parameter view. We design a universal checkpoint format—incorporating merged parameter representations and mapping metadata—a Universal Checkpoint Language (UCL), and an on-demand state reconstruction mechanism, achieving, for the first time, full decoupling of checkpointing from parallel configurations. Evaluated on LLaMA, Bloom, and other mainstream large models under diverse parallelism paradigms—including tensor parallelism (TP), pipeline parallelism (PP), data parallelism (DP), and context parallelism (CP)—our approach reduces post-failure recovery time by 12–28% on average, significantly enhancing cross-configuration portability and system robustness.
To address communication-intensive session programming, this paper proposes a novel session type system supporting explicit commit, rollback, and abort operations. To prevent illegal cross-participant state restoration, it introduces—within the session types framework—the first statically decidable rollback compliance check, ensuring rollbacks affect only locally accessed states and never violate inter-participant boundaries. Building on session type theory, the authors extend the session language and integrate MAUDE for design-time, type-level verification. They formally prove that the system satisfies error-freedom and progress properties. The core contributions are: (i) a safe and controllable rollback semantics grounded in session types; (ii) static enforcement of cross-participant isolation during rollback; and (iii) verifiable session recovery behavior, enabling rigorous reasoning about fault-tolerant distributed protocols.
Training large foundation models (LFMs) faces significant challenges in checkpoint management, including poor cross-framework compatibility, tight coupling with parallelization strategies, heterogeneous storage backends, and severe I/O bottlenecks. To address these, this work proposes an industrial-grade unified archival system. Its core contributions are: (1) a novel parallelism-agnostic checkpoint serialization format; (2) a full-stack I/O optimization framework integrating a dynamic resharding engine, multi-framework abstraction interfaces (PyTorch/Megatron/DeepSpeed), asynchronous high-throughput storage adapters, and a distributed I/O monitoring toolchain; and (3) runtime support for cross-parallelism resharding, multi-backend adaptivity, and rapid failure recovery. Experiments demonstrate an average 54.20× reduction in checkpoint blocking time, with peak checkpoint save and load speedups of 9.96× and 8.80×, respectively. The system has been stably deployed in production environments scaling to over one thousand GPUs.
Manually implementing checkpoint/restart mechanisms for MPI-based scientific applications is time-consuming and heavily reliant on expert knowledge. This work proposes an automated framework that leverages the Claude Opus 4.7 large language model in conjunction with OpenCode CLI to iteratively generate code, achieving end-to-end fault tolerance injection without human intervention for the first time. Evaluated on six real-world MPI applications, the approach automatically produces valid checkpointing code within an average of 50 minutes, incurs negligible runtime overhead, and achieves recovery performance comparable to handcrafted implementations. This significantly lowers the technical barrier to resilience engineering in high-performance computing applications.
Current quantum high-performance computing systems lack effective fault-tolerance and recovery mechanisms, and conventional checkpointing approaches based on quantum state preservation are fundamentally hindered by the no-cloning theorem. This work proposes a novel algorithm-level fault-tolerance framework that reframes checkpointing and recovery as problems of control flow and algorithmic state management, thereby avoiding direct storage of quantum states. Instead, it leverages mid-circuit measurements, classical feedforward, and conditional operations in dynamic quantum circuits to efficiently capture and restore program execution states. The approach enables reliable interruption and resumption for representative iterative or staged quantum algorithms—including variational eigensolvers, the Quantum Approximate Optimization Algorithm (QAOA), and time-stepping simulations—significantly enhancing the reliability and restartability of quantum computational tasks.
研究了AI代理系统中检查点和回滚(C/R)的安全性问题,通过建立执行模型识别出五种基本故障模式,并展示了它们对安全的影响。
This study addresses the high implementation barrier and reliance on domain expertise associated with checkpoint/restart mechanisms in HPC scientific applications by proposing an automated fault-tolerance framework driven by large language model (LLM) coding agents. The framework establishes a generate-verify-revise closed-loop pipeline that autonomously injects fault-tolerance capabilities into MPI applications without human intervention. This work provides the first demonstration that LLM agents can efficiently perform resilience engineering given state visibility. Experimental results show that the proposed approach successfully produces 41 functional implementations, each requiring under one hour on average, while introducing negligible runtime overhead. Furthermore, the achieved fault recovery efficiency is comparable to that of manually written code, substantially reducing development costs.
This study addresses the miscalculation of recovery costs in checkpointing systems for coding agents, which stems from the inability to distinguish valid safety boundaries. We propose a counterfactual checkpoint advantage framework that drives dual branches toward identical logical failures and matches conditional recoveries to quantify the performance differential between saving and skipping states. This approach reveals that recovery fundamentally entails re-derivation rather than simple replay, demonstrating that traditional evaluation metrics based on preserved workload introduce severe bias. Controlled experiments on the SWE-bench Verified benchmark show that this strategy saves an average of 49.4 seconds across twelve tasks. Furthermore, the observed first-checkpoint transition rate is 0.64, whereas subsequent checkpoints yield negative values, thereby establishing a performance baseline for optimal checkpoint placement.
LayerCheck通过选择性保存更新超过阈值的层来解决大语言模型训练中的检查点效率问题,减少了I/O开销和恢复时间。
This work addresses the inefficiency of existing fault-tolerance mechanisms for large language model (LLM) agents, which rely on full restarts or application-level checkpointing and struggle to recover critical GPU state—such as KV caches and scheduler metadata—efficiently. The authors propose a device-resident, persistent kernel implementation that enables fine-grained, low-overhead checkpointing and recovery directly within the GPU’s native execution context, thereby avoiding CPU bottlenecks. By injecting incremental checkpoint logic into the PTX/SASS layer via JIT compilation, and combining GPU module load interception, lock-free ring buffers, and CXL/host-memory logging, the approach achieves framework-agnostic, transparent fault tolerance. It supports efficient dirty-page detection and incremental snapshots of key structures like KV caches and adapter pages, substantially reducing work loss and recovery latency upon failure.