🤖 AI Summary
This study addresses the challenges of privacy leakage and non-independent and identically distributed (Non-IID) data in federated deep clustering by proposing the Fed-BRDECS framework. Methodologically, it introduces a local stability loss to replace the transmission of global statistics, thereby preserving privacy, and incorporates prediction-balanced sampling alongside a centroid restart mechanism to effectively mitigate class imbalance without requiring labels. By integrating federated learning, deep embedded clustering, and adaptive sampling techniques, the proposed framework significantly outperforms mainstream baselines across image, text, and time-series anomaly detection benchmarks. It achieves improved clustering accuracy without increasing inference costs, realizing a dual breakthrough in both privacy protection and adaptation to heterogeneous data distributions.
📝 Abstract
Federated deep clustering seeks to learn clustering-friendly representations from decentralized unlabeled data while preserving client privacy. However, Deep Embedded Clustering (DEC)-style objectives depend on global soft-assignment statistics that require clients to reveal their sensitive information. We propose Fed-BRDECS, a privacy-preserving and heterogeneity-aware federated deep embedded clustering framework. Fed-BRDECS replaces the globally normalized clustering objective with a locally computable sample-stability loss, avoiding the transmission of local soft-assignment distributions. To tackle non-IID client distributions, we introduce prediction-balanced sampling, which oversamples locally rare predicted clusters without requiring ground-truth labels, and centroid-level restarting, which periodically refreshes biased or inactive centroids. Experiments on image and text clustering benchmarks show that Fed-BRDECS consistently outperforms representative federated clustering and deep clustering baselines under both IID and non-IID partitions. We further demonstrate its applicability to federated time-series anomaly detection, where it improves reconstruction-based detectors without adding inference-time cost.