🤖 AI Summary
This study addresses the bottleneck that selective state space models (SSMs) lack architecture-aware convergence theory in federated learning by establishing, for the first time, a convergence analysis framework tailored to modern selective SSMs such as Mamba2. By deriving gradient and smoothness bounds for single- and multi-layer SSMs, this work rigorously establishes convergence upper bounds for FedAvg and FedProx, revealing the critical impact of recurrence stability on distributed optimization. Cross-domain experiments validate the theoretical bounds and demonstrate the explanatory power of SSM-specific limits for practical federated learning behavior. Ultimately, this research overcomes the limitations of conventional approaches that overlook state parameterization, providing a principled theoretical foundation for deploying selective SSMs in federated settings.
📝 Abstract
Modern state space models (SSMs), such as Mamba2, provide a compelling alternative to transformers by combining linear-time sequence modeling with recurrent state-space dynamics. However, the behavior of SSMs in distributed learning settings remains poorly understood. In particular, the existing standard federated learning methods are largely architecture-agnostic, and do not account for the stability, selectivity, and state-space parameterization that characterize modern selective SSMs. To address this, we derive architecture-aware gradient and smoothness bounds for single- and multi-layer selective SSMs, and convergence bounds for FedAvg and FedProx, characterizing how recurrent stability, input-dependent discretization, and state projection norms affect federated optimization. We then numerically validate the single-layer bounds on sequences generated by a teacher SSM, using a learner that follows the analyzed recurrence. We use this analysis to formulate expectations about the effects of local training and client heterogeneity, and examine these expectations by comparing nine federated learning algorithms on Mamba2 language modeling across six text domains. These experiments illustrate how SSM-specific bounds can provide a basis for interpreting the behavior of practical federated learning algorithms.