🤖 AI Summary
This study addresses the lack of theoretical guarantees for the stability and convergence of federated learning under large step-size settings with heterogeneous devices. Focusing on linearly separable data, this work integrates multiclass logistic regression with convex optimization theory to analyze the dynamics of FedAvg under arbitrarily large step sizes and heterogeneous local updates, corroborated by numerical experiments. It rigorously establishes that the algorithm remains stable and convergent even with large step sizes, revealing that the impact of device heterogeneity vanishes asymptotically. Furthermore, the objective value converges to zero at an O(1/R) rate and decreases monotonically in later stages, bounded by the average number of local steps. This research fills a critical theoretical gap in federated optimization under extreme parameter configurations.
📝 Abstract
This paper revisits the distributed learning problem for training a multinomial logistic regression model with the Federated Averaging ($\texttt{FedAvg}$) algorithm. We concentrate on a scenario with arbitrarily large stepsizes and heterogeneous update rules where the devices may perform a different number of local updates in each round. We show that, with linearly separable data, $\texttt{FedAvg}$ is stable with any stepsizes and the objective values converge to zero at the rate of ${\cal O}(1/R)$, where $R$ is the number of communication rounds. Our result also demonstrates that the effects of device heterogeneity vanish asymptotically. For sufficiently large $R$, the objective values decrease monotonically and is bounded by ${\cal O}( 1 / (R T_{\rm avg}))$, where $T_{\rm avg}$ is the average number of local update steps per communication round across devices. Numerical experiments support our findings.