ZOCheck: CPU-Shadow Checkpointing for Zeroth-Order LLM Fine-Tuning

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
针对零阶优化在大模型微调中的容错问题,ZOCheck通过CPU影子进程持续回放日志更新,并异步保存恢复镜像,实现高效故障恢复。
📝 Abstract
Zeroth-order (ZO) optimization is an attractive option for memory-efficient LLM fine-tuning, but its fault tolerance remains underexplored. Unlike first-order training, ZO progress can be represented by lightweight seed-and-scalar step logs, yet naive log-only recovery still incurs replay cost that grows with training progress, and shortcut replay does not preserve the executed floating-point trajectory. We present ZOCheck, a fault-tolerant ZO training system that exploits this replayable structure through a CPU shadow process that continuously replays logged updates, materializes consistent recovery images off the GPU critical path, and persists them asynchronously. ZOCheck therefore combines non-blocking checkpointing during training with fast recovery from a near-current state. We also develop a cost model for choosing the snapshot policy under realistic failure rates. Experiments show that ZOCheck reduces checkpoint overhead by up to 219.7x and recovery latency by 1.55x on average compared with asynchronous full-state checkpointing, translating into up to 21.3x lower end-to-end wasted time across the evaluated failure rates, while preserving exact recovery behavior.
Problem

Research questions and friction points this paper is trying to address.

Zeroth-Order Optimization
Fault Tolerance
LLM Fine-Tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

CPU-Shadow Checkpointing
Zeroth-Order Optimization
Non-blocking Checkpointing
Fast Recovery