Score
Designs and implements techniques that map model predictions, internal activations, or input features to topological summaries computed by topological data analysis—particularly persistent homology outputs such as persistence diagrams, barcodes, and landscapes—to visualize and quantify how models use geometric and connectivity patterns. Builds analyses, metrics, and visualizations that highlight persistent structural features, attribute feature importance via homological persistence, and produce interpretable explanations of model decisions based on multiscale topological structure.
Persistent homology offers multiscale topological robustness, yet existing statistical methods can only detect global topological differences without localizing their sources. To address this, we propose a novel framework integrating graphical models with Bayesian inference: each persistent bar’s birth and death times are modeled as events on a graph (e.g., MST edge additions or cycle formations), and a conic-space latent variable is introduced to construct an interpretable probabilistic model. Using an exponential likelihood and hierarchical latent structure, our method enables scalable, cross-group Bayesian inference. This is the first approach enabling precise localization and mechanistic interpretation of topological discrepancies. Applied to Alzheimer’s disease neuroimaging data, it successfully identifies topologically aberrant brain regions with biological interpretability. The implementation is open-source and designed for extension to high-dimensional, complex datasets.
In supervised learning based on persistent homology, computing persistence diagrams is computationally expensive, and conventional full matrix reduction often discards essential topological information from the original data. To address this, we propose a novel paradigm that directly extracts topological feature vectors from the **unreduced boundary matrix**, bypassing costly reduction while preserving richer algebraic topological structure. Our method is grounded in persistent homology theory and introduces a differentiable, scalable feature mapping mechanism. We conduct systematic evaluations across diverse datasets and tasks—including classification and regression. Experiments demonstrate that our approach matches or surpasses standard reduced-persistence baselines in predictive performance, while substantially reducing computational complexity. These results empirically validate our core claim: strong discriminative topological features can be obtained *without full matrix reduction*. The work thus establishes a new pathway toward efficient topological machine learning.
Conventional entropy-based diversity metrics fail to capture the intrinsic geometric and topological structure of high-dimensional training data, limiting their ability to assess data quality and predict model performance. Method: This work introduces persistent homology—a tool from topological data analysis—to quantify structural properties such as connected components, holes, and higher-order voids, thereby jointly characterizing data richness and redundancy through both topological and geometric lenses. Contribution/Results: Experiments demonstrate that the proposed topological metrics exhibit strong correlation with model generalization performance and serve as effective predictors of data quality. The framework enables principled data selection and efficient training, offering a novel paradigm for dataset curation. By leveraging topological signatures, it enhances training efficiency and robustness of AI systems without requiring model retraining or architectural modification. The approach is broadly applicable across domains where data geometry and topology critically influence learning dynamics.
This study investigates the internal mechanisms underlying the transition from memorization to generalization—known as “grokking”—in neural network training. Focusing on modular arithmetic tasks, the authors analyze the topological evolution of model embedding point clouds using persistent homology, complemented by Fourier analysis and local intrinsic dimensionality estimates, to systematically compare representational structures under different data regimes. They report the first evidence that grokking coincides with a pronounced increase in both the maximum and total persistence of first-order homology groups, and demonstrate that this topological shift correlates directly with generalization rather than mere memorization. These findings reveal how the cyclic structure inherent in the task is geometrically and topologically encoded in the representation space, offering a unified perspective on generalization in deep learning.
In neuroimaging research, the diversity of rfMRI brain representations—e.g., parcellation schemes and feature types—induces inconsistent individual variability estimates, severely undermining result reproducibility and cross-study integration. Method: We propose a topological similarity assessment framework based on persistent homology, specifically designed for small-sample, high-dimensional neuroimaging data. It integrates Vietoris–Rips complexes, topological bootstrapping, and hierarchical clustering, and introduces a novel prevalence-weighted Wasserstein distance to enable unbiased topological comparison across samples and heterogeneous metric spaces. Contribution/Results: We demonstrate that low-persistence but high-prevalence homological generators encode interpretable biological signals. Validation on large-scale cohorts reveals that environmental dimensions of representations—not decomposition rank—predominantly govern topological feature count and stability. The framework enables robust, unbiased clustering across diverse representations, advancing reproducible, integrative neuroimaging analysis.
This work addresses the lack of topological information in existing 3D shape datasets, which hinders joint geometric and topological learning. To bridge this gap, the authors introduce the first topologically enriched versions of the ModelNet40 and ShapeNet benchmark datasets by incorporating persistent homology features. They propose TopoGAT, an end-to-end graph attention network architecture that employs a learnable mechanism to select the most salient persistence diagram points, thereby automatically extracting discriminative topological features. Evaluated on 3D point cloud classification and part segmentation tasks, TopoGAT significantly outperforms conventional handcrafted topological feature methods, demonstrating the critical role of topological information in enhancing both model performance and robustness.
Traditional persistence diagrams struggle to capture the interactive topological relationships between point clouds and lack cross-structural modeling capacity. This work establishes, for the first time, the existence and statistical foundations of cross-persistence diagram densities and introduces an end-to-end framework that integrates topological data analysis, statistical learning, and deep learning to directly predict these densities from point cloud coordinates and distance matrices. A novel noise-augmentation mechanism is innovatively incorporated to enhance the discriminative power of point clouds, significantly extending the applicability of topological data analysis in cross-structural settings. Experiments demonstrate that the proposed method achieves state-of-the-art performance in both density prediction and point cloud discrimination across multiple datasets, while also showing promising potential in geometric analyses of time series and AI-generated text.
This work addresses the insufficient modeling of topological structures in existing deep learning approaches for 3D point clouds, where persistent homology has largely been relegated to peripheral roles. The authors introduce 3DPHDL—the first systematic design space that deeply integrates persistent homology as a structural inductive bias throughout the entire point cloud learning pipeline. This integration encompasses six well-defined injection points spanning simplicial complex construction, filtration strategies, persistence representations, and their coordination with backbone architectures. Through controlled experiments on PointNet, DGCNN, and Point Transformer—augmented with persistence diagrams, images, and landscapes—on ModelNet40 and ShapeNetPart, the approach significantly improves accuracy in classification and segmentation, enhances part consistency, and boosts robustness to noise and sampling variations, while also revealing inherent trade-offs between representational capacity and computational complexity.
This study addresses the effectiveness of topological feature extraction for univariate time series classification by mapping time series into graph structures using five complex network methods, including visibility graphs, transition graphs, and proximity graphs. Persistent diagrams are generated via Vietoris–Rips filtration and persistent homology, then vectorized using persistence landscapes and topological statistics. The work reveals that both the choice of network construction and distance metric critically influence classification performance: diffusion distance consistently outperforms shortest-path distance, and optimal graph representations vary across signal types. Furthermore, the robustness of topological features under noise is empirically validated. Experiments on twelve UCR benchmark datasets demonstrate that while no single network construction universally dominates, diffusion distance consistently yields superior results.
该研究提出一种在Mapper诱导的结构化表示上进行学习的框架,解决传统方法可能忽略数据多尺度结构的问题。