🤖 AI Summary
This study investigates the intrinsic relationship between graph visualization and clustering algorithms in preserving multidimensional data structures. Centered on the Fruchterman-Reingold (FR) force-directed layout combined with PCA-based dimensionality reduction, we systematically quantify structural similarity differences among four agglomerative hierarchical clustering methods: single linkage, complete linkage, average linkage, and Ward’s method. Results indicate that while these clustering approaches exhibit high mutual consistency, they all deviate significantly from the original data structure. In contrast, FR visualization more faithfully preserves the topological characteristics of high-dimensional data. To our knowledge, this work is the first to reveal, from a quantitative perspective, the fundamental distinctions between force-directed layouts and hierarchical clustering, thereby providing empirical evidence for the synergistic selection of visualization and clustering techniques in high-dimensional data exploration.
📝 Abstract
Graph visualization methods and agglomerative clustering have been frequently considered in data analysis and pattern recognition. Because these approaches are interrelated and complementary, it is of particular interest to investigate their associations. In this work, we study the possible relationship between the Fruchterman-Reingold graph visualization method and four types of agglomerative clustering adopting single- and complete-linkage, average, and Ward's linkage criteria. Three types of datasets have been considered in 2 and 10 dimensions, as well as the PCA projection of the latter to two dimensions. The results obtained suggest that the relationship between the methods considered did not vary much for the three types of data mentioned above. At the same time, the agglomerative methods tended to yield results that are mostly similar to each other, while presenting moderate similarity with the original data. The Fruchterman-Reingold visualization resulted similar to the original data, but exhibited relatively smaller similarity to the agglomerative methods.