Provably Convergent Decentralized Optimization over Directed Graphs under Generalized Smoothness

📅 2026-01-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of the classical Lipschitz smoothness assumption in scenarios involving rapidly varying gradients and highly heterogeneous data by proposing a decentralized optimization algorithm tailored for directed communication graphs. The method integrates gradient tracking with adaptive gradient clipping and, within the generalized $(L_0, L_1)$-smoothness framework, establishes the first convergence guarantee that does not rely on the bounded gradient dissimilarity assumption. Experimental evaluations on LIBSVM and CIFAR-10 datasets—using regularized logistic regression and convolutional neural networks, respectively—demonstrate that the proposed algorithm consistently outperforms existing approaches in both convergence speed and stability.

Technology Category

Machine Learning: OptimizationSearch and Optimization: Non-convex OptimizationConstraint Satisfaction and Optimization: Distributed CSP/Optimization

Application Category

Graph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsSystems and Infrastructure for Web, Mobile and WoT: Experiences and lessons learnt from Web-based algorithms and system deploymentsResponsible Web: Human-perceived consequences of algorithmic deployment on the web
📝 Abstract
Decentralized optimization has become a fundamental tool for large-scale learning systems; however, most existing methods rely on the classical Lipschitz smoothness assumption, which is often violated in problems with rapidly varying gradients. Motivated by this limitation, we study decentralized optimization under the generalized $(L_0, L_1)$-smoothness framework, in which the Hessian norm is allowed to grow linearly with the gradient norm, thereby accommodating rapidly varying gradients beyond classical Lipschitz smoothness. We integrate gradient-tracking techniques with gradient clipping and carefully design the clipping threshold to ensure accurate convergence over directed communication graphs under generalized smoothness. In contrast to existing distributed optimization results under generalized smoothness that require a bounded gradient dissimilarity assumption, our results remain valid even when the gradient dissimilarity is unbounded, making the proposed framework more applicable to realistic heterogeneous data environments. We validate our approach via numerical experiments on standard benchmark datasets, including LIBSVM and CIFAR-10, using regularized logistic regression and convolutional neural networks, demonstrating superior stability and faster convergence over existing methods.
Problem

Research questions and friction points this paper is trying to address.

decentralized optimization
generalized smoothness
directed graphs
gradient dissimilarity
non-Lipschitz
Innovation

Methods, ideas, or system contributions that make the work stand out.

generalized smoothness
decentralized optimization
directed graphs
gradient clipping
gradient tracking
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yanan Bo
Department of Electrical and Computer Engineering, Clemson University, Clemson, SC 29634, USA
Yongqiang Wang
Yongqiang Wang
Clemson University
cooperative controldistributed optimizationprivacy