An Efficient Compression of Deep Neural Network Checkpoints Based on Prediction and Context Modeling

📅 2025-06-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the high storage overhead and accuracy degradation during checkpoint (model weights + optimizer states) recovery in deep neural network training, this paper proposes a prediction-driven context modeling compression method. Our approach leverages prior checkpoints as predictive sources to guide arithmetic coding’s context modeling—a novel design first introduced in this work—and jointly optimizes compression ratio and training recovery fidelity through co-designed pruning, quantization, and predictive coding. Experiments demonstrate that our method achieves an average 3.2× compression ratio, significantly reducing checkpoint bitrates, while maintaining near-lossless recovery: post-recovery Top-1 accuracy degradation remains below 0.1%. This work establishes a new paradigm for efficient distributed training and fault-tolerant recovery in storage-constrained environments.

Technology Category

Machine Learning: Learning on the Edge & Model CompressionData Mining & Knowledge Management: Data CompressionCognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingGraph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphs
📝 Abstract
This paper is dedicated to an efficient compression of weights and optimizer states (called checkpoints) obtained at different stages during a neural network training process. First, we propose a prediction-based compression approach, where values from the previously saved checkpoint are used for context modeling in arithmetic coding. Second, in order to enhance the compression performance, we also propose to apply pruning and quantization of the checkpoint values. Experimental results show that our approach achieves substantial bit size reduction, while enabling near-lossless training recovery from restored checkpoints, preserving the model's performance and making it suitable for storage-limited environments.
Problem

Research questions and friction points this paper is trying to address.

Efficient compression of neural network checkpoints
Reducing bit size while preserving model performance
Enabling near-lossless training recovery from compressed checkpoints
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prediction-based compression using arithmetic coding
Pruning and quantization for enhanced compression
Near-lossless training recovery from compressed checkpoints
💼 Related Jobs
No related jobs found.
Y
Yuriy Kim
ITMO University, Saint-Petersburg, Russia
Evgeny Belyaev
Evgeny Belyaev
ITMO University
Video coding