Score
Designs and implements methods that compute and summarize topological characteristics of data—such as persistence diagrams, Betti curves, and Hurewicz-type invariants—and extract those summaries as feature descriptors. Builds and evaluates pipelines and models that encode and fuse these topological features into vector representations and integrate them with other data representations for topology-informed learning and analysis.
This work addresses the ongoing challenge of effectively incorporating topological priors into optimization problems. It proposes a systematic framework for topological optimization grounded in persistent homology, which unifies gradient-based optimization with a differentiable topological regularization loss to enable end-to-end learning of topological features. Providing a comprehensive survey of theoretical and algorithmic advances over the past decade, the paper offers—for the first time—an accessible, unified introduction tailored to mathematicians and data scientists new to the field. Accompanied by an open-source library, this contribution aims to lower entry barriers and foster broader adoption of topological data analysis in machine learning and data science communities.
In supervised learning based on persistent homology, computing persistence diagrams is computationally expensive, and conventional full matrix reduction often discards essential topological information from the original data. To address this, we propose a novel paradigm that directly extracts topological feature vectors from the **unreduced boundary matrix**, bypassing costly reduction while preserving richer algebraic topological structure. Our method is grounded in persistent homology theory and introduces a differentiable, scalable feature mapping mechanism. We conduct systematic evaluations across diverse datasets and tasks—including classification and regression. Experiments demonstrate that our approach matches or surpasses standard reduced-persistence baselines in predictive performance, while substantially reducing computational complexity. These results empirically validate our core claim: strong discriminative topological features can be obtained *without full matrix reduction*. The work thus establishes a new pathway toward efficient topological machine learning.
Persistent diagrams (PDs) generated by persistent homology reside in non-Hilbert spaces, hindering their direct integration into machine learning pipelines. To address this, we propose a unified, robust, and interpretable vectorization framework for PDs. Our open-source toolkit—efficiently implemented in R and Python via Rcpp, NumPy, and Cython—integrates multiple established methods, including persistence images, persistence landscapes, and Betti curves, while supporting customizable kernel functions, normalization schemes, and grid parameters. It incorporates rigorous mathematical definitions and best-practice guidelines. The framework is natively compatible with scikit-learn and tidymodels, enabling batch vectorization through kernel-based embeddings, functional integration, and statistical aggregation. Empirical evaluation across multiple benchmark datasets demonstrates that the resulting Euclidean vectors preserve topological discriminative power, achieve 3–10× faster inference, and substantially lower the engineering barrier to incorporating topological data analysis (TDA) into standard ML workflows.
This work addresses the modeling bottleneck of non-Euclidean data—such as point clouds, graphs, and meshes—by proposing a topology-driven machine learning paradigm. Methodologically, it introduces the first systematic three-tier framework: “topological representation → topological embedding → topological guidance.” It employs the Euler Characteristic Transform (ECT) as a differentiable topological invariant to enable end-to-end embedding of persistent homology features. By integrating topological priors with geometric deep learning, it designs topology-aware modules and hybrid neural architectures. Contributions include: (1) significantly improved robustness and efficiency in non-Euclidean data modeling; (2) provision of structured inductive biases and novel tools for interpretable analysis in neural networks; and (3) establishment of foundational theory and a computationally tractable framework for topological machine learning.
Conventional entropy-based diversity metrics fail to capture the intrinsic geometric and topological structure of high-dimensional training data, limiting their ability to assess data quality and predict model performance. Method: This work introduces persistent homology—a tool from topological data analysis—to quantify structural properties such as connected components, holes, and higher-order voids, thereby jointly characterizing data richness and redundancy through both topological and geometric lenses. Contribution/Results: Experiments demonstrate that the proposed topological metrics exhibit strong correlation with model generalization performance and serve as effective predictors of data quality. The framework enables principled data selection and efficient training, offering a novel paradigm for dataset curation. By leveraging topological signatures, it enhances training efficiency and robustness of AI systems without requiring model retraining or architectural modification. The approach is broadly applicable across domains where data geometry and topology critically influence learning dynamics.
Commercial and economic data often exhibit nonlinear, multiscale structures that linear methods fail to capture effectively. To address this, we introduce topological data analysis (TDA), constructing simplicial complexes via persistent homology and integrating SAX/eSAX symbolic representation with multiscale distance metrics to extract robust topological features. Our key contribution is the Topological Stability Index (TSI), a novel interpretable metric quantifying structural variability and providing actionable insights into systemic fluctuations. We validate the framework across three real-world domains—consumer behavior, stock market dynamics, and foreign exchange time series—demonstrating its reproducible ability to uncover clustering structures and temporal patterns missed by conventional statistical approaches. Results show that TDA substantially enhances pattern discovery in complex business data, while TSI exhibits strong discriminative stability and domain-relevant interpretability.
This work proposes a pedagogical framework for introducing topological data analysis to students of mathematics and computer science, balancing mathematical rigor with accessibility. Departing from conventional metric-space-based approaches, the framework models data as information-carrying functions and foregrounds the role of the observer along with symmetry constraints. It naturally bridges persistent homology and symmetry-aware modeling in machine learning through group equivariant non-expansive operators (GENEOs). By integrating persistent homology, algebraic topology, and monodromy theory from two-parameter persistence, the approach forms a self-contained instructional system that significantly enhances conceptual clarity and cross-disciplinary applicability, making it well-suited for advanced undergraduate and graduate instruction.
This work addresses the sensitivity of the Mapper algorithm to lens functions, cover parameters, and clustering strategies, for which no systematic evaluation framework previously existed. The authors propose the first triaxial assessment framework that comprehensively evaluates Mapper variants across three complementary dimensions: stability, cluster quality, and topological shape preservation. Experiments on synthetic data and the UCI handwritten digits dataset reveal inherent trade-offs among these dimensions, demonstrating that no single configuration achieves optimal performance across all metrics simultaneously. The study further identifies a “topological explosion” phenomenon at high resolutions, offering practical guidance for parameter selection in real-world applications and highlighting key challenges for future research in Mapper-based topological data analysis.
This work addresses a critical limitation in existing vectorization and kernel methods for topological data analysis, which neglect the rank-induced inclusion relations among persistence intervals, thereby distorting structural information and reducing interpretability. To overcome this, the authors propose high-order persistence diagrams that explicitly capture structural dependencies by recursively modeling interval inclusion relationships. They further introduce, for the first time, harmonic analysis and the zeta transform to implicitly aggregate high-order diagrams in the spectral domain, reducing computational complexity from quadratic to nearly linear. Experimental results demonstrate that the proposed method significantly outperforms explicit aggregation strategies on random network models, achieving superior efficiency and scalability while preserving structural fidelity.
Traditional persistence diagrams struggle to capture the interactive topological relationships between point clouds and lack cross-structural modeling capacity. This work establishes, for the first time, the existence and statistical foundations of cross-persistence diagram densities and introduces an end-to-end framework that integrates topological data analysis, statistical learning, and deep learning to directly predict these densities from point cloud coordinates and distance matrices. A novel noise-augmentation mechanism is innovatively incorporated to enhance the discriminative power of point clouds, significantly extending the applicability of topological data analysis in cross-structural settings. Experiments demonstrate that the proposed method achieves state-of-the-art performance in both density prediction and point cloud discrimination across multiple datasets, while also showing promising potential in geometric analyses of time series and AI-generated text.
本文通过构建一个包含广泛Z-不变量的数据集,并使用神经网络从中提取拓扑信息,如同调类和基础图结构,探索了低维拓扑中的模式识别问题。