🤖 AI Summary
This work addresses the excessive memory overhead of adaptive optimizers like Adam in edge-based federated learning, where storing high-precision momentum and variance states hinders the deployment of large models and multitask scenarios. The study is the first to reveal the distinct statistical distributions of momentum and variance in Federated Adam and leverages this insight to propose a distribution-aware 8-bit asymmetric quantization scheme: momentum is compressed via block-wise linear encoding, while variance employs logarithmic-space encoding. Crucially, model parameters remain in full precision. Experiments on CIFAR-10/100 demonstrate a 3.37× reduction in optimizer memory footprint with no accuracy loss under moderate data heterogeneity, and up to a 5.74 percentage point accuracy gain (p<0.01) under extreme heterogeneity, enabling efficient federated learning on resource-constrained devices.
📝 Abstract
Federated learning on edge devices must cope with non-IID client data and tight memory budgets. Adaptive optimizers like Adam stabilize training under data heterogeneity but require storing full-precision momentum and variance states, often tripling client memory overhead. This limits deployable model sizes and concurrent federated jobs on resource-constrained devices.
We empirically observe that momentum and variance in federated Adam exhibit fundamentally different statistical properties: momentum values are symmetric and bounded, while variance spans eight orders of magnitude with log-normal structure. Motivated by this asymmetry, we propose \textbf{Q-LocalAdam}, which applies distribution-aware 8-bit quantization block-wise linear encoding for momentum and log-space encoding for variance while keeping model parameters in full precision.
Across CIFAR-10 and CIFAR-100 under varying data heterogeneity ($α\in \{0.1, 0.5, 1.0, \text{IID}\}$), Q-LocalAdam achieves $3.37\times$ optimizer memory reduction with no accuracy loss under moderate heterogeneity and significant improvements under extreme heterogeneity (e.g., +5.74pp on CIFAR-100, $α=0.1$). Multi-seed validation confirms statistical significance ($p<0.01$). In contrast, naive uniform quantization degrades to random performance, demonstrating that distribution-aware design is essential. Q-LocalAdam enables larger models and more concurrent workloads on memory-constrained edge devices without modifying the federated protocol.