🤖 AI Summary
Persistent diagrams (PDs) generated by persistent homology reside in non-Hilbert spaces, hindering their direct integration into machine learning pipelines. To address this, we propose a unified, robust, and interpretable vectorization framework for PDs. Our open-source toolkit—efficiently implemented in R and Python via Rcpp, NumPy, and Cython—integrates multiple established methods, including persistence images, persistence landscapes, and Betti curves, while supporting customizable kernel functions, normalization schemes, and grid parameters. It incorporates rigorous mathematical definitions and best-practice guidelines. The framework is natively compatible with scikit-learn and tidymodels, enabling batch vectorization through kernel-based embeddings, functional integration, and statistical aggregation. Empirical evaluation across multiple benchmark datasets demonstrates that the resulting Euclidean vectors preserve topological discriminative power, achieve 3–10× faster inference, and substantially lower the engineering barrier to incorporating topological data analysis (TDA) into standard ML workflows.
📝 Abstract
Persistent homology is a widely-used tool in topological data analysis (TDA) for understanding the underlying shape of complex data. By constructing a filtration of simplicial complexes from data points, it captures topological features such as connected components, loops, and voids across multiple scales. These features are encoded in persistence diagrams (PDs), which provide a concise summary of the data's topological structure. However, the non-Hilbert nature of the space of PDs poses challenges for their direct use in machine learning applications. To address this, kernel methods and vectorization techniques have been developed to transform PDs into machine-learning-compatible formats. In this paper, we introduce a new software package designed to streamline the vectorization of PDs, offering an intuitive workflow and advanced functionalities. We demonstrate the necessity of the package through practical examples and provide a detailed discussion on its contributions to applied TDA. Definitions of all vectorization summaries used in the package are included in the appendix.