🤖 AI Summary
This study addresses the challenge of applying conventional statistical methods to collections of heterogeneous networks that vary in size and type and lack node correspondence. To overcome this, the authors propose a functional Topological Data Analysis (funTDA) framework that uniquely integrates functional data analysis with persistent homology to extract topological features from networks. This approach enables standard statistical operations—including mean and variance estimation, principal component analysis, and hypothesis testing—despite the non-Euclidean nature of network structures, thereby establishing a unified inferential framework. Empirical evaluations demonstrate that funTDA effectively discriminates networks with distinct connectivity patterns and successfully uncovers significant topological differences in real-world applications, such as literary co-occurrence networks and influenza gene regulatory networks.
📝 Abstract
Statistical analysis of collections of networks, where each network is treated as the primary unit of observation, is of growing importance across a wide range of application domains, including gene regulatory, social, and financial networks. As networks consist of vertices and edges that do not naturally reside in Euclidean space, the direct application of conventional statistical methodologies, such as the computation of means and covariances, principal component analysis, and hypothesis testing, to samples of networks is not straightforward. A central challenge lies in defining meaningful measures of similarity or distance between networks of potentially varying sizes and structural types (e.g., directed, undirected, weighted or unweighted), particularly when no predefined node correspondence exists. To address these challenges, we introduce a framework termed functional topological data analysis (funTDA), which integrates tools from functional data analysis and topological data analysis to facilitate exploratory data analysis and inference on samples of networks. The proposed framework enables the computation of summary statistics, including means and variances, and supports the application of principal component analysis and hypothesis testing to topological features extracted from network data. Through simulation studies involving networks with varying connectivity structures, we demonstrate the ability of funTDA to distinguish between distinct network configurations. The methodology is illustrated through two real-data applications: networks constructed from pairwise word co-occurrences in novels by Jane Austen and Charles Dickens, and gene regulatory networks derived from gene expression measurements for seventeen individuals exposed to H3N2 influenza. In both applications, differences in network topology are assessed using principal component analysis and hypothesis testing.