🤖 AI Summary
This study addresses the absence of a unified theoretical framework for network partitioning clustering that accounts for both hard/soft assignment mechanisms and non-vertex-centered cluster prototypes. The authors systematically analyze four classical models: the hard-assignment p-median problem (PMP) and spectral clustering (SSC), alongside the soft-assignment probabilistic density clustering (PDC) and fuzzy c-means (FCM). They reveal, for the first time, that optimal solutions of PMP and PDC are inherently confined to graph vertices, whereas SSC and FCM can yield cluster centers located along edges. Through rigorous mathematical analysis of their structural properties and optimization behaviors under network topology, the work elucidates the critical roles of bottleneck points and vertex-constrained solutions, thereby establishing a theoretical foundation for efficient clustering algorithms in facility location, network design, and similarity search via graph embeddings.
📝 Abstract
This study presents a theoretical analysis of partitional clustering on networks, analyzing both hard and soft assignment schemes with different objective functions. Cluster centers are not restricted to vertices but can also be located along the edges. We examine four key models: P-Median (PMP) and Sum of Squares Clustering (SSC) under hard assignment, and Probabilistic Distance Clustering (PDC) and Fuzzy C-Means (FCM) under soft assignment. Through mathematical analysis, we uncover structural properties that differentiate these models, such as the significance of assignment bottleneck points and the role of vertex-restricted solutions in determining optimal cluster centers. Our findings reveal that, while SSC and FCM can yield optimal centers along edges, PMP and PDC inherently favor vertex placement, leading to insights into clustering behavior on networks. These insights offer new directions for designing efficient algorithms and have implications ranging from facility location and network design to clustering on the embedding graphs that power similarity search in modern retrieval systems.