build federated learning systems

Designs and implements end-to-end federated learning systems that coordinate distributed model training across many client devices, including centralized and fully decentralized/peer-to-peer protocols, aggregation algorithms (such as federated averaging and cluster-based aggregation), simulators, and deployment-grade implementations that handle non‑IID data, client churn, mobility and node failures. Builds and evaluates privacy-preserving and secure aggregation methods, and trains or analyzes application models (for example behavioral monitoring, behavioral fingerprinting, intrusion detection, and object detection) without pooling raw data, while measuring protocol performance, robustness, privacy guarantees, and accuracy.

buildfederatedlearningsystems

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.52
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$195K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Federated learning (FL) confronts fundamental challenges including statistical heterogeneity, privacy preservation, and collaborative efficiency. To address these, this work establishes a hybrid research framework integrating bibliometric analysis and systematic review. It proposes the first multi-level taxonomy of aggregation techniques, structured along three core dimensions: personalization, optimization, and robustness. Through empirical evaluation, it systematically benchmarks mainstream FL architectures and synchronization strategies under both IID and non-IID data settings. Furthermore, it develops a reproducible benchmarking platform for aggregation methods. The study identifies critical technical bottlenecks and provides theoretically grounded guidance and practical pathways for emerging directions—including privacy-enhancing aggregation, heterogeneity-aware modeling, and robust aggregation. Collectively, this work significantly advances the rigor, reproducibility, and extensibility of systematic FL research.

Addressing data privacy, heterogeneity, and efficiency in decentralized AIProviding taxonomy and experiments for future FL research directionsSurveying federated learning aggregation techniques and challenges

This work addresses the security and scalability challenges in federated learning arising from adversarial gradient updates and aggregation bottlenecks by proposing the first end-to-end distributed architecture integrating zero-knowledge proofs (ZKPs). By compiling machine learning loss functions into Rank-1 Constraint Systems (R1CS), the approach enables cryptographic verification of local computations at each node without accessing raw gradients, thereby effectively mitigating model poisoning attacks. Experimental results demonstrate that the system maintains high throughput even at a scale of one thousand nodes and achieves a model accuracy retention rate of 94.2% under adversarial conditions, marking the first scalable federated learning framework that simultaneously guarantees strong security and high performance.

Distributed AIFederated LearningModel Poisoning

This work addresses privacy preservation in training machine learning models on edge devices using users’ private local data. We propose a Private Federated Learning (PFL) framework tailored for mobile app selection tasks. Methodologically, we design a lightweight neural architecture integrating attention mechanisms and uncertainty modeling to dynamically capture evolving user behavior on-device; further, we employ a differential privacy-enhanced aggregation strategy to enable collaborative model training without exposing raw user data. To the best of our knowledge, this is the first end-to-end deployment and empirical validation of PFL in real-world mobile environments. Experimental results demonstrate that the model achieves continuously improving accuracy as user behavior evolves, while strictly complying with GDPR and other privacy regulations. The framework successfully balances strong privacy guarantees, practical utility, and efficient adaptability to resource-constrained edge devices.

Improving app selection model accuracyPrivate Federated Learning on edge devicesTraining model with user data privacy

This work identifies a critical oversight in federated learning research: while existing studies emphasize network topology, they neglect the fundamental distinction between centralized (CFL) and decentralized federated learning (DFL)—namely, their training protocols (decoupled aggregation vs. joint optimization). To address this, we propose the first protocol-centric taxonomy for CFL/DFL. We systematically expose a long-standing research gap concerning distributed optimization methods in DFL and establish a theoretical triadic trade-off model among privacy, robustness, and model utility. Leveraging protocol-driven analysis, distributed optimization theory, and meta-review methodology, we develop a unified analytical framework for CFL and DFL. Our framework rigorously characterizes how distributed optimization fundamentally enables DFL, thereby providing a principled foundation and design guidance for next-generation federated learning systems that are secure, scalable, and adversarially robust.

Analyzes impact of protocols on model utility, privacy, and attack robustness.Explores differences between centralized and decentralized Federated Learning protocols.Identifies lack of research on decentralized FL using distributed optimization methods.

Enhancing Federated Learning Through Secure Cluster-Weighted Client Aggregation

Mar 29, 2025
KR
Kanishka Ranaweera
🏛️ Deakin University | The University of Queensland

In federated learning, client data heterogeneity degrades model performance, impairs convergence stability, and exacerbates privacy risks. To address these challenges, we propose ClusterGuardFL—a dynamic weighted aggregation framework. Its core contributions are threefold: (1) adaptive clustering based on a dissimilarity score to identify semantically similar clients; (2) cluster-size-aware weighting to mitigate bias from small clusters; and (3) point-level reconciliation guided by confidence-aware softmax weights, jointly enhancing robustness, fairness, and differential privacy compatibility. The framework integrates k-means clustering, model dissimilarity measurement, confidence modeling, and secure aggregation. Extensive experiments on multi-source heterogeneous datasets demonstrate that ClusterGuardFL significantly improves global model accuracy and convergence stability, effectively suppressing interference from malicious or low-quality clients. Results validate the synergistic benefits of its weighted aggregation strategy in simultaneously strengthening robustness and privacy preservation.

Addressing data heterogeneity in federated learning systemsEnsuring privacy and fairness across diverse user groupsImproving model performance and convergence in FL

Latest Papers

What's happening recently
View more

This work addresses key limitations in secure aggregation for federated learning—namely, excessive communication rounds, high computational overhead from public-key operations, and poor robustness to client dropouts—by introducing a secret sharing–based distributed aggregator architecture. In this approach, a small committee of clients acts as aggregators: each participant secret-shares its local model update among committee members, who then compute partial aggregation results locally and return shares that enable the server to efficiently reconstruct the global model. By eliminating conventional local masking and homomorphic encryption, the proposed method substantially reduces both computation and communication costs. Experimental results demonstrate that, under a realistic setting with 100,000-dimensional update vectors and 100,000 5G clients, the protocol achieves a 4.6× speedup over the OPA protocol while maintaining strong privacy guarantees and system efficiency.

Client DropoutsCommunication EfficiencyFederated Learning

This work addresses the vulnerability of asynchronous federated learning to malicious aggregators, which can compromise model integrity and client data privacy, thereby threatening system liveness and confidentiality. To counter this, the paper proposes the first asynchronous secure federated learning framework resilient to Byzantine aggregators. The approach leverages a replicated aggregator architecture, decoupled secure aggregation, and differential privacy via Gaussian noise, effectively mitigating Byzantine attacks without requiring consensus among aggregators. Additionally, a participation-balancing strategy is introduced to dynamically harmonize privacy budgets and model bias in asynchronous settings. Experimental results demonstrate that the proposed method maintains competitive training performance while simultaneously ensuring strong privacy guarantees, system liveness, and robustness against adversarial aggregators.

asynchronous federated learningByzantine aggregatorsclient privacy

This work addresses the limited robustness of federated learning under extreme conditions characterized by non-independent and identically distributed (Non-IID) client data and a majority (>50%) of malicious participants. To this end, the authors propose a heuristic defense algorithm that integrates server-side learning, client update filtering, and geometric median-based aggregation. Notably, the method operates effectively even when the server possesses only a small amount of real or synthetic data whose distribution significantly diverges from that of the clients—a setting previously unaddressed in the literature. Experimental results demonstrate that the proposed approach substantially improves model accuracy under such highly adversarial scenarios, thereby confirming its strong robustness and practical efficacy.

Federated LearningMalicious AttacksNon-IID Data

This work proposes FedPLT, a novel federated learning approach designed to address the high communication and computational overhead, strong device heterogeneity, and issues of inconsistent parameter distributions and biased global loss estimation caused by existing partial-parameter training methods. FedPLT employs a structured partial-layer training strategy that adaptively assigns each client a personalized subset of the model based on its resource capacity. By integrating resource-aware model partitioning, hierarchical parameter selection, optimal client sampling, and aggregation optimization, FedPLT achieves performance on par with or superior to FedAvg while using only 18%–29% of trainable parameters. The method significantly reduces the number of straggler clients and demonstrates superior performance in highly heterogeneous environments compared to current state-of-the-art approaches.

Communication OverheadComputation OverheadDevice Heterogeneity

Hot Scholars

XL

Xunkai Li

School of Computer Science and Technology, Beijing Institution of Technology
Data-centric AIGraph MLAI4Science
CG

Christopher G. Brinton

Elmore Associate Professor of ECE, Purdue University
NetworkingMachine LearningCommunicationsEdge Computing
MF

Minghong Fang

University of Louisville
SecurityPrivacyAI SafetyMachine Learning
WN

Wei Ni

FIEEE, AAIA Fellow, Senior Principal Scientist & Conjoint Professor, CSIRO/UNSW
6G security and privacyconnected and trusted intelligenceapplied AI/ML
ZX

Ziyue Xu

NVIDIA
Medical Image AnalysisComputer VisionFederated Learning