🤖 AI Summary
This study addresses the limitation of relying on a single uncertainty metric to simultaneously detect heterogeneous data corruptions, such as image noise and label flipping, in federated learning. Employing a ResNet-20 architecture with Monte Carlo Dropout under Dirichlet non-IID data partitioning, this work compares the detection performance—evaluated via AUC—of predictive loss and entropy-based uncertainty. It proposes an adaptive strategy that matches detection signals to specific corruption types. The findings reveal a complementary mechanism: predictive loss is more sensitive to label flipping (AUC > 0.85), whereas entropy uncertainty proves superior for detecting image noise. By overcoming the constraints of single-metric uncertainty estimation, this research establishes an adaptive data quality assessment framework tailored for federated learning environments.
📝 Abstract
Federated learning (FL) data corruption can affect either inputs or labels, but it remains unclear whether input-conditional uncertainty and prediction-label loss expose these corruption modes equally. This paper compares two corruption-detection signals in FL: input-conditional uncertainty and prediction-label loss. The uncertainty signal is characterised using a learned aleatoric variance estimate together with Monte Carlo (MC) dropout variance and entropy measures, while the loss is computed against the supplied label. We test these signals against additive image noise and persistent random label flips. On ResNet-20 with CIFAR-10 and SVHN under Dirichlet partitions with data that are not independent and identically distributed (non-IID), the two corruption types behave differently. For persistent random label flips, the within-client per-sample area under the receiver operating characteristic curve (AUC) is 0.85 on CIFAR-10 and 0.95 on SVHN for prediction-label loss, while every uncertainty estimator stays at chance (0.49--0.50). This pattern is consistent with the model remaining confident in the underlying image despite the supplied label being wrong. For image noise, expected-entropy uncertainty rises above chance (0.67 on CIFAR-10 and 0.66 on SVHN), while loss responds comparably (0.64 on both). Each signal is therefore the stronger detector for a different corruption: the prediction-label loss for persistent label flips, and expected-entropy uncertainty for image noise, with its advantage becoming apparent as federation-wide corruption prevalence increases. Robust FL data-quality assessment should match the signal to the corruption rather than rely on uncertainty alone across corruption types.