Comparing representations of high-dimensional data with persistent homology: a case study in neuroimaging

📅 2023-06-23
📈 Citations: 1
Influential: 1
📄 PDF

career value

219K/year
🤖 AI Summary
In neuroimaging research, the diversity of rfMRI brain representations—e.g., parcellation schemes and feature types—induces inconsistent individual variability estimates, severely undermining result reproducibility and cross-study integration. Method: We propose a topological similarity assessment framework based on persistent homology, specifically designed for small-sample, high-dimensional neuroimaging data. It integrates Vietoris–Rips complexes, topological bootstrapping, and hierarchical clustering, and introduces a novel prevalence-weighted Wasserstein distance to enable unbiased topological comparison across samples and heterogeneous metric spaces. Contribution/Results: We demonstrate that low-persistence but high-prevalence homological generators encode interpretable biological signals. Validation on large-scale cohorts reveals that environmental dimensions of representations—not decomposition rank—predominantly govern topological feature count and stability. The framework enables robust, unbiased clustering across diverse representations, advancing reproducible, integrative neuroimaging analysis.
📝 Abstract
Despite much attention, the comparison of reduced-dimension representations of high-dimensional data remains a challenging problem in multiple fields, especially when representations remain high-dimensional compared to sample size. We offer a framework for evaluating the topological similarity of high-dimensional representations of very high-dimensional data, a regime where topological structure is more likely captured in the distribution of topological"noise"than a few prominent generators. Treating each representational map as a metric embedding, we compute the Vietoris-Rips persistence of its image. We then use the topological bootstrap to analyze the re-sampling stability of each representation, assigning a"prevalence score"for each nontrivial basis element of its persistence module. Finally, we compare the persistent homology of representations using a prevalence-weighted variant of the Wasserstein distance. Notably, our method is able to compare representations derived from different samples of the same distribution and, in particular, is not restricted to comparisons of graphs on the same vertex set. In addition, representations need not lie in the same metric space. We apply this analysis to a cross-sectional sample of representations of functional neuroimaging data in a large cohort and hierarchically cluster under the prevalence-weighted Wasserstein. We find that the ambient dimension of a representation is a stronger predictor of the number and stability of topological features than its decomposition rank. Our findings suggest that important topological information lies in repeatable, low-persistence homology generators, whose distributions capture important and interpretable differences between high-dimensional data representations.
Problem

Research questions and friction points this paper is trying to address.

Comparing inter-subject variability across fMRI brain representations
Assessing reproducibility challenges in brain-behavior association studies
Evaluating feature type impact on neuroimaging results comparability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Topological data analysis for brain variability
Persistent homology compares 34 representations
Feature type critical for replicability assessment
🔎 Similar Papers
No similar papers found.
T
T. Easley
Mallinckrodt Institute of Radiology, Washington University in St. Louis
K
Kevin Freese
IBM
E
E. Munch
Dept of Computational Mathematics, Science and Engineering, Dept of Mathematics, Michigan State University
J
J. Bijsterbosch
Mallinckrodt Institute of Radiology, Washington University School of Medicine