Score
Designs and implements methods that turn data into simplicial complexes or Mapper covers, compute persistent homology and persistence diagrams, and produce vectorized or summary representations of those topological invariants. Builds analysis and ML components that use these summaries — e.g., distance comparisons, clustering, artifact detection, sampling-aware computations, and persistence-based featureization or regularization losses — to analyze structure in point-clouds, similarity networks, or filtration-indexed datasets.
This work addresses the sensitivity of the Mapper algorithm to lens functions, cover parameters, and clustering strategies, for which no systematic evaluation framework previously existed. The authors propose the first triaxial assessment framework that comprehensively evaluates Mapper variants across three complementary dimensions: stability, cluster quality, and topological shape preservation. Experiments on synthetic data and the UCI handwritten digits dataset reveal inherent trade-offs among these dimensions, demonstrating that no single configuration achieves optimal performance across all metrics simultaneously. The study further identifies a “topological explosion” phenomenon at high resolutions, offering practical guidance for parameter selection in real-world applications and highlighting key challenges for future research in Mapper-based topological data analysis.
Analyzing and comparing non-hierarchical multi-scale clustering sequences remains challenging due to the lack of stable, topology-aware representations. Method: We propose Multi-scale Clustering Filtration (MCF), a stable simplicial filtration that encodes clustering partitions at arbitrary scales. We systematically introduce persistent homology to this task by constructing MCF and its equivalent nerve complex, proving that in the hierarchical case it reduces to the Vietoris–Rips filtration on an ultrametric space. Contribution/Results: Empirical validation on synthetic data demonstrates that zero- and higher-dimensional persistence diagrams derived from MCF serve as robust topological features, effectively characterizing and distinguishing diverse multi-scale clustering structures. MCF thus establishes a novel paradigm for the quantitative evaluation and comparative analysis of multi-scale clustering, enabling principled, topology-driven assessment beyond traditional metrics.
This work addresses the challenge in Mapper algorithms where the cover parameter requires manual tuning and fails to adapt to intrinsic data structure. We propose a data-adaptive, automated cover optimization method. Our core innovations are: (i) the first integration of G-means clustering with the Anderson–Darling normality test to statistically determine cover interval boundaries; and (ii) the incorporation of Gaussian Mixture Models (GMMs) to guide semantically informed cover splitting. Evaluated on both synthetic and real-world datasets, our method significantly improves structural fidelity and semantic interpretability of Mapper graphs. Moreover, it achieves an order-of-magnitude speedup over iterative baseline approaches. The implementation is publicly available.
In neuroimaging research, the diversity of rfMRI brain representations—e.g., parcellation schemes and feature types—induces inconsistent individual variability estimates, severely undermining result reproducibility and cross-study integration. Method: We propose a topological similarity assessment framework based on persistent homology, specifically designed for small-sample, high-dimensional neuroimaging data. It integrates Vietoris–Rips complexes, topological bootstrapping, and hierarchical clustering, and introduces a novel prevalence-weighted Wasserstein distance to enable unbiased topological comparison across samples and heterogeneous metric spaces. Contribution/Results: We demonstrate that low-persistence but high-prevalence homological generators encode interpretable biological signals. Validation on large-scale cohorts reveals that environmental dimensions of representations—not decomposition rank—predominantly govern topological feature count and stability. The framework enables robust, unbiased clustering across diverse representations, advancing reproducible, integrative neuroimaging analysis.
To address robust clustering of unstructured point clouds, this paper proposes a novel spectral clustering method based on the global topological contribution of point pairs. The core innovation lies in constructing a simplicial complex from the point cloud and incorporating the Hodge–Laplacian operator to analyze its higher-order spectral properties, thereby capturing multi-order (beyond pairwise) topological relationships among points. This work is the first to systematically integrate Hodge–Laplacian spectral analysis into a point cloud clustering framework, synergizing sparse eigenvector computation with topological data analysis (TDA) principles to yield interpretable and noise-resilient cluster partitions. Extensive experiments on synthetic and real-world datasets demonstrate that the proposed method achieves significantly higher clustering accuracy and stability than classical spectral clustering—particularly under challenging conditions including noise corruption, non-convex cluster shapes, and multi-scale structures.
Conventional entropy-based diversity metrics fail to capture the intrinsic geometric and topological structure of high-dimensional training data, limiting their ability to assess data quality and predict model performance. Method: This work introduces persistent homology—a tool from topological data analysis—to quantify structural properties such as connected components, holes, and higher-order voids, thereby jointly characterizing data richness and redundancy through both topological and geometric lenses. Contribution/Results: Experiments demonstrate that the proposed topological metrics exhibit strong correlation with model generalization performance and serve as effective predictors of data quality. The framework enables principled data selection and efficient training, offering a novel paradigm for dataset curation. By leveraging topological signatures, it enhances training efficiency and robustness of AI systems without requiring model retraining or architectural modification. The approach is broadly applicable across domains where data geometry and topology critically influence learning dynamics.
Persistent homology computation on point clouds suffers from high computational complexity and difficulty in simultaneously preserving both global and local topological fidelity. To address this, we propose the δ-core subsampling method—the first to integrate strong collapse theory into topology-aware subsampling. Our approach constructs a minimal core point set satisfying a δ-neighborhood condition, ensuring strong collapse equivalence between the original point cloud and the subsample, thereby provably preserving all persistent homology groups. Integrated with strong-collapse-driven simplicial complex simplification and persistent homology analysis, the method robustly retains salient topological features across multiple scales. Experiments on synthetic and real-world datasets demonstrate that our method achieves an average 12.7% improvement in persistence approximation accuracy over state-of-the-art subsampling strategies, while accelerating computation by 3.2–5.8×.
Persistent Laplacians (PL) lack stable, finite-dimensional vector representations, hindering their practical use in machine learning. Method: We propose the first spectral-feature-driven vectorization framework for PL, introducing the Persistent Laplacian Diagram (PLD) and Persistent Laplacian Image (PLI)—constructing finite-dimensional vector embeddings via spectral discretization and image-based representation grounded in PL’s spectral theory. We design real-valued signature functions and rigorously prove that PLI is Lipschitz stable under input perturbations. Results: Experiments demonstrate that PLD/PLI distinguish graph structures indiscernible to both standard persistent homology and combinatorial Laplacians, while preserving richer geometric and combinatorial information. This significantly enhances topological data representation capability. Our work establishes the first vectorization paradigm for PL with theoretical guarantees—namely, stability and discriminability—and practical applicability.
This work addresses the insufficient modeling of topological structures in existing deep learning approaches for 3D point clouds, where persistent homology has largely been relegated to peripheral roles. The authors introduce 3DPHDL—the first systematic design space that deeply integrates persistent homology as a structural inductive bias throughout the entire point cloud learning pipeline. This integration encompasses six well-defined injection points spanning simplicial complex construction, filtration strategies, persistence representations, and their coordination with backbone architectures. Through controlled experiments on PointNet, DGCNN, and Point Transformer—augmented with persistence diagrams, images, and landscapes—on ModelNet40 and ShapeNetPart, the approach significantly improves accuracy in classification and segmentation, enhances part consistency, and boosts robustness to noise and sampling variations, while also revealing inherent trade-offs between representational capacity and computational complexity.