Robustifying Asynchronous SGD via Soft Throttling

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of asynchronous stochastic gradient descent (SGD) to Byzantine attacks in distributed learning, where fast clients disproportionately dominate model updates. To mitigate this, we propose Throttle, an algorithm that unifies asynchronous and synchronous Byzantine-robust SGD frameworks. Throttle employs a dynamic weighting mechanism based on gradient arrival times, exponentially down-weighting fast clients to defend against malicious nodes. Theoretical analysis establishes the convergence rate of the proposed algorithm. Empirical evaluations demonstrate that Throttle significantly enhances system robustness under Byzantine attacks, effectively balancing the efficiency of asynchronous training with the security guarantees of synchronous methods. Notably, it also outperforms standard asynchronous SGD in benign, attack-free settings.
📝 Abstract
Asynchronous SGD is a popular algorithm for distributed learning where each client's gradient update is applied on arrival. This leads to a speed-up, but also an increased vulnerability to attacks, as fast clients can dominate the total update. We introduce Throttle, a Byzantine-robust generalization of asynchronous SGD where the key idea is to exponentially down-weight updates from faster clients by a factor $q$. Both asynchronous SGD ($q=1$) and synchronous Byzantine-robust SGD ($q\to\infty$) correspond to specific settings of Throttle. We provide a theoretical analysis of the convergence rate and validate the robustness to attacks both theoretically and empirically. Remarkably, our experiments show that this down-weighting mechanism can also improve performance over standard asynchronous SGD even in the non-Byzantine setting.
Problem

Research questions and friction points this paper is trying to address.

Asynchronous SGD
Distributed Learning
Byzantine Robustness
Byzantine Attacks
🔎 Similar Papers
No similar papers found.