🤖 AI Summary
This paper addresses distributed clustering under partial feature space overlap, where multiple parties hold private, heterogeneous datasets with partially overlapping features (e.g., cross-institutional healthcare data) and cannot share raw data or full feature sets. To tackle this, we propose two novel federated clustering algorithms: (i) a global centroid aggregation scheme that enables federated updates of shared cluster centers, and (ii) a statistical modeling approach that generates and aggregates synthetic proxy data to align heterogeneous feature spaces. Both methods support participant autonomy in selecting local clustering models and customizing computational overhead. Under mild regularity conditions, the algorithms converge to the centralized optimal solution. Experiments on three public benchmark datasets demonstrate that our methods achieve clustering performance close to the centralized oracle baseline and significantly outperform existing distributed clustering baselines, confirming their practical deployability.
📝 Abstract
We introduce and address a novel distributed clustering problem where each participant has a private dataset containing only a subset of all available features, and some features are included in multiple datasets. This scenario occurs in many real-world applications, such as in healthcare, where different institutions have complementary data on similar patients. We propose two different algorithms suitable for solving distributed clustering problems that exhibit this type of feature space heterogeneity. The first is a federated algorithm in which participants collaboratively update a set of global centroids. The second is a one-shot algorithm in which participants share a statistical parametrization of their local clusters with the central server, who generates and merges synthetic proxy datasets. In both cases, participants perform local clustering using algorithms of their choice, which provides flexibility and personalized computational costs. Pretending that local datasets result from splitting and masking an initial centralized dataset, we identify some conditions under which the proposed algorithms are expected to converge to the optimal centralized solution. Finally, we test the practical performance of the algorithms on three public datasets.