Distributed clustering in partially overlapping feature spaces

📅 2025-10-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses distributed clustering under partial feature space overlap, where multiple parties hold private, heterogeneous datasets with partially overlapping features (e.g., cross-institutional healthcare data) and cannot share raw data or full feature sets. To tackle this, we propose two novel federated clustering algorithms: (i) a global centroid aggregation scheme that enables federated updates of shared cluster centers, and (ii) a statistical modeling approach that generates and aggregates synthetic proxy data to align heterogeneous feature spaces. Both methods support participant autonomy in selecting local clustering models and customizing computational overhead. Under mild regularity conditions, the algorithms converge to the centralized optimal solution. Experiments on three public benchmark datasets demonstrate that our methods achieve clustering performance close to the centralized oracle baseline and significantly outperform existing distributed clustering baselines, confirming their practical deployability.

Technology Category

Machine Learning: ClusteringConstraint Satisfaction and Optimization: Distributed CSP/OptimizationSearch and Optimization: Distributed Search

Application Category

User Modeling, Personalization and Recommendation: Federated recommendation systems and personalizationSystems and Infrastructure for Web, Mobile and WoT: Federated Web and WoT systems, including distributed, federated and edge-based data processingGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphs
📝 Abstract
We introduce and address a novel distributed clustering problem where each participant has a private dataset containing only a subset of all available features, and some features are included in multiple datasets. This scenario occurs in many real-world applications, such as in healthcare, where different institutions have complementary data on similar patients. We propose two different algorithms suitable for solving distributed clustering problems that exhibit this type of feature space heterogeneity. The first is a federated algorithm in which participants collaboratively update a set of global centroids. The second is a one-shot algorithm in which participants share a statistical parametrization of their local clusters with the central server, who generates and merges synthetic proxy datasets. In both cases, participants perform local clustering using algorithms of their choice, which provides flexibility and personalized computational costs. Pretending that local datasets result from splitting and masking an initial centralized dataset, we identify some conditions under which the proposed algorithms are expected to converge to the optimal centralized solution. Finally, we test the practical performance of the algorithms on three public datasets.
Problem

Research questions and friction points this paper is trying to address.

Clustering data distributed across multiple institutions with overlapping features
Solving feature space heterogeneity when each site has partial feature sets
Developing federated and one-shot algorithms for distributed clustering scenarios
Innovation

Methods, ideas, or system contributions that make the work stand out.

Federated algorithm collaboratively updates global centroids
One-shot algorithm merges synthetic proxy datasets centrally
Local clustering flexibility with personalized computational costs
🔎 Similar Papers
No similar papers found.
A
Alessio Maritan
University of Padova
L
Luca Schenato
University of Padova