🤖 AI Summary
This work addresses the critical limitation of existing serverless federated learning systems, which struggle to train large models due to strict memory constraints imposed by serverless functions. To overcome this barrier, the authors propose GradsSharding, a novel approach that shards gradient tensors and processes them in parallel across serverless functions, with each function aggregating only its assigned shard. This method achieves mathematically equivalent results to conventional tree-based aggregation while attaining constant memory consumption per function—decoupled from the number of participating clients—for the first time. Implemented on AWS Lambda, the sharded FedAvg algorithm demonstrates effectiveness across model sizes ranging from 43 MB to 5 GB, reduces training costs by 2.7× on VGG-16, and stands as the only serverless federated learning solution capable of surpassing the 10 GB memory ceiling.
📝 Abstract
Federated learning (FL) aggregation on serverless platforms faces a hard scalability ceiling: existing architectures (lambda-FL, LIFL) partition clients across aggregators, but every aggregator must hold the complete model gradient in memory. When gradients exceed the per-function memory limit (e.g., 10 GB on AWS Lambda), aggregation becomes infeasible regardless of tree depth or branching factor. We propose GradsSharding, which instead partitions the gradient tensor into M shards, each averaged independently by a serverless function that receives contributions from all clients. Because FedAvg averaging is element-wise, this produces bit-identical results to tree-based approaches, so model accuracy is invariant by construction. Per-function memory is bounded at O(|θ|/M), independent of client count, enabling aggregation of arbitrarily large models. We evaluate GradsSharding against lambda-FL and LIFL through HPC experiments and real AWS Lambda deployments across model sizes from 43 MB to 5 GB. Results show a cost crossover at approximately 500 MB gradient size, 2.7x cost reduction at VGG-16 scale, and that GradsSharding is the only architecture that remains deployable beyond the serverless memory ceiling.