ADMM Algorithms for Residual Network Training: Convergence Analysis and Parallel Implementation

📅 2023-10-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses gradient explosion during backpropagation in residual network training and the high memory consumption and communication overhead inherent in distributed training. We propose a serial and parallel optimization framework based on proximal (linearized) ADMM. Theoretically, we establish, for the first time without assumptions on network width, depth, or dataset size, the R-linear convergence of the algorithm. Practically, we design a low-communication, low-memory distributed coordination protocol enabling localized parameter updates. Experiments demonstrate rapid and stable convergence, improved test accuracy, and significantly reduced per-node memory usage; the parallel variant substantially enhances scalability and training efficiency. Our core contribution lies in systematically integrating ADMM into residual network training—achieving both rigorous theoretical guarantees and practical engineering viability.
📝 Abstract
We propose both serial and parallel proximal (linearized) alternating direction method of multipliers (ADMM) algorithms for training residual neural networks. In contrast to backpropagation-based approaches, our methods inherently mitigate the exploding gradient issue and are well-suited for parallel and distributed training through regional updates. Theoretically, we prove that the proposed algorithms converge at an R-linear (sublinear) rate for both the iteration points and the objective function values. These results hold without imposing stringent constraints on network width, depth, or training data size. Furthermore, we theoretically analyze our parallel/distributed ADMM algorithms, highlighting their reduced time complexity and lower per-node memory consumption. To facilitate practical deployment, we develop a control protocol for parallel ADMM implementation using Python's multiprocessing and interprocess communication. Experimental results validate the proposed ADMM algorithms, demonstrating rapid and stable convergence, improved performance, and high computational efficiency. Finally, we highlight the improved scalability and efficiency achieved by our parallel ADMM training strategy.
Problem

Research questions and friction points this paper is trying to address.

Mitigating exploding gradient in residual network training
Enabling parallel and distributed training via ADMM
Proving R-linear convergence without strict network constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proximal ADMM for residual network training
Parallel implementation reduces time complexity
Control protocol for multiprocessing deployment
🔎 Similar Papers
2024-06-30arXiv.orgCitations: 1
💼 Related Jobs
No related jobs found.
Tsinghua University | Alibaba Group
J
Jintao Xu
Department of Mathematical Sciences, Tsinghua University, Beijing 100084, China
Y
Yifei Li
Alibaba Group (at the time of this work, Independent)
W
Wenxun Xing
Department of Mathematical Sciences, Tsinghua University, Beijing 100084, China