🤖 AI Summary
This work addresses the lack of systematic and reproducible benchmarks for evaluating robustness in federated learning across diverse datasets and model architectures. It introduces the first comprehensive evaluation matrix comprising 500 experimental configurations, spanning five aggregation methods, five datasets, five model architectures, and four attack types—including sign-flipping, Gaussian noise, and BadNets backdoor attacks. Through rigorous reproduction and log-based auditing, the study systematically compares method performance under both clean and adversarial conditions. Results show that Trimmed Mean achieves the highest average accuracy (76.02%) in clean settings, while Krum demonstrates superior robustness against specific attacks. The analysis also uncovers critical implementation flaws—such as the misuse of the TTLR metric and inconsistencies between prediction and aggregation updates in FedPARETO—thereby significantly enhancing the transparency and auditability of robustness evaluations in federated learning.
📝 Abstract
Robust comparisons of federated aggregation methods require joint consideration of predictive performance, threat definitions, metric semantics, and execution provenance. A 500-cell seed-1 evaluation matrix was reconstructed across five aggregation methods, five datasets, five architectures, and four recorded conditions: clean, sign-flipping, Gaussian, and BadNets. Successful execution logs were identified for 454 original runs and 36 repaired or rerun executions, whereas 10 clean SVHN cells were supported by summary-only provenance. Trimmed Mean achieved the highest clean macro-mean accuracy (76.02%) and the lowest mean within-task rank (1.70). Krum attained the highest recorded accuracy under both sign-flipping and Gaussian configurations. These relative rankings remained unchanged when analysis was restricted to 21 task pairs for which original successful logs were available for every method-condition combination. Audit of the supplied BadNets metric implementation established that every test input is triggered prior to target-label counting; consequently, the retained metric represents Triggered Target-Label Rate (TTLR) rather than a conventional target-excluding attack success rate. An audit of the supplied FedPARETO scaffold further identified a pathway in which predictive summaries may characterize an uncorrupted local model while the aggregation weight is applied to a separately corrupted update, introducing a potential discrepancy between reported predictive outcomes and the updates used for aggregation. The canonical matrix contains a single identified seed for each cell, and exact attack and configuration lineage is incomplete. Accordingly, the findings should be interpreted as descriptive comparisons within the recorded configurations and not as statistical estimates or universal claims regarding robustness.