Fed-BRDECS: Privacy-Preserving and Heterogeneity-Aware Federated Deep Embedded Clustering

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of privacy leakage and non-independent and identically distributed (Non-IID) data in federated deep clustering by proposing the Fed-BRDECS framework. Methodologically, it introduces a local stability loss to replace the transmission of global statistics, thereby preserving privacy, and incorporates prediction-balanced sampling alongside a centroid restart mechanism to effectively mitigate class imbalance without requiring labels. By integrating federated learning, deep embedded clustering, and adaptive sampling techniques, the proposed framework significantly outperforms mainstream baselines across image, text, and time-series anomaly detection benchmarks. It achieves improved clustering accuracy without increasing inference costs, realizing a dual breakthrough in both privacy protection and adaptation to heterogeneous data distributions.
📝 Abstract
Federated deep clustering seeks to learn clustering-friendly representations from decentralized unlabeled data while preserving client privacy. However, Deep Embedded Clustering (DEC)-style objectives depend on global soft-assignment statistics that require clients to reveal their sensitive information. We propose Fed-BRDECS, a privacy-preserving and heterogeneity-aware federated deep embedded clustering framework. Fed-BRDECS replaces the globally normalized clustering objective with a locally computable sample-stability loss, avoiding the transmission of local soft-assignment distributions. To tackle non-IID client distributions, we introduce prediction-balanced sampling, which oversamples locally rare predicted clusters without requiring ground-truth labels, and centroid-level restarting, which periodically refreshes biased or inactive centroids. Experiments on image and text clustering benchmarks show that Fed-BRDECS consistently outperforms representative federated clustering and deep clustering baselines under both IID and non-IID partitions. We further demonstrate its applicability to federated time-series anomaly detection, where it improves reconstruction-based detectors without adding inference-time cost.
Problem

Research questions and friction points this paper is trying to address.

Federated deep clustering
Privacy preservation
Deep Embedded Clustering
Non-IID data
Data heterogeneity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Federated Deep Embedded Clustering
Privacy-Preserving
Sample-Stability Loss
Prediction-Balanced Sampling
Centroid-Level Restarting
🔎 Similar Papers
No similar papers found.
H
Haemin Park
Northwestern University
Diego Klabjan
Diego Klabjan
Northwestern University
Machine learning
M
Martin W. Braun
Intel Corporation
X
Xiuqi Li
Intel Corporation
B
Balakrishnan Ananthanarayanan
Intel Corporation