topological data analysis

Extracting and using multi-scale topological or geometric summaries to represent feature evolution, capture forms of ill-posedness, and serve as surrogates for unobserved concept changes in models.

topologicaldataanalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Traditional machine learning struggles to effectively model shape data with nonlinear geometric structures and their intrinsic variability. This work proposes a unified analytical framework that systematically integrates differential geometry, manifold statistics, and geometric deep learning to address the challenges posed by complex, unaligned shapes exhibiting nonlinear variation. The framework encompasses key components including shape representation, geodesic metrics, parametrization, and statistical inference. It has been successfully applied to multiscale biological geometric data—such as cellular morphologies and primate dental evolution—revealing structural patterns and evolutionary trajectories underlying shape variation. This approach establishes both a theoretical foundation and a practical paradigm for geometry-aware learning in shape analysis.

geometric datageometric variationmachine learning

Must-Read Papers

Most classic and influential ideas
View more

Cover Learning for Large-Scale Topology Representation

Mar 12, 2025
LS
Luis Scoccola
🏛️ Centre de Recherches Mathématiques | Institut des sciences mathématiques | Université du Québec à Montréal | Université de Sherbrooke | Queen Mary University of London | Max Planck Institute of Molecular Cell Biology and Genetics | Technische Universität Dresden | University of Oxford

This work addresses the challenge of topological representation for large-scale geometric data, where existing methods—such as Mapper and geometric simplicial complexes—suffer from deteriorating performance at scale, heavy reliance on manual hyperparameter tuning, and limited expressivity in capturing high-dimensional topological features. We propose the first differentiable optimization framework that learns topology-preserving covers end-to-end, eliminating dependence on predefined geometric structures or hand-crafted design. Our core innovation is modeling the cover as a differentiable neural network and introducing topologically grounded loss functions—e.g., nerve consistency—to enable gradient-based parameterization and optimization of covering subsets. The resulting simplicial complexes are more compact yet exhibit superior preservation of high-dimensional topology. Extensive evaluation demonstrates significant improvements over traditional methods on large datasets, with enhanced scalability, robustness to data scale, and fidelity to underlying topological structure.

Learning topologically-faithful covers for geometric datasets.Optimizing cover learning to improve large-scale topology representation.Overcoming limitations of geometric complexes and Mapper graphs.

Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling

Oct 01, 2025
FT
Federico Tiblias
🏛️ Technical University of Darmstadt

Prior work designs geometric structures for specific features, lacking generalizability across model families and scales. Method: We propose Supervised Multidimensional Scaling (SMDS), the first method to automatically discover feature manifolds across diverse large language models (LLMs) and sizes, systematically characterizing the geometric representations of concepts in activation space. By integrating concept-labeled data into manifold learning, SMDS uncovers distinct geometric structures—circular, linear, and cluster-like—in temporal reasoning, corresponding respectively to periodicity, sequentiality, and categorical semantics. Results: We empirically demonstrate that these structures dynamically reconfigure within context and functionally support concretized reasoning. Experiments confirm their cross-model stability, interpretability, and strong correlation with downstream reasoning performance.

Analyzing how manifolds support reasoning and context adaptationAutomatically discovering feature manifolds in language modelsRevealing geometric structures of concepts in latent spaces

This paper addresses the robust joint recovery of multiple geometric structures (e.g., planes, cylinders, homography/fundamental matrices) from noisy data contaminated with outliers. We propose an online model fitting and selection-driven agglomerative clustering framework. Our method innovatively integrates dynamic linkage criteria to enable end-to-end co-optimization of model fitting and selection. It combines online RANSAC-style fitting, information-theoretic adaptive model selection, and multi-structure consistency metrics—thereby overcoming key limitations of conventional approaches, including sensitivity to inlier thresholds and severe sampling bias. Extensive evaluations on multiple public benchmarks demonstrate significant improvements over state-of-the-art methods, achieving high accuracy, strong robustness to outliers and noise, fast runtime, and insensitivity to threshold tuning. The source code is publicly available.

Recover multiple geometric structures from noisy dataRobust fitting for mixed parametric model classesSimultaneous handling of multi-class model recovery

Existing diversity metrics for datasets predominantly rely on statistical distributions or entropy, often overlooking the intrinsic geometric structure of the data. This work introduces persistence landscapes (PLs)—a tool from topological data analysis—into diversity assessment, offering a geometric perspective to quantify structural diversity and establishing a direct link between geometric features and diversity. The proposed PLDiv metric is grounded in rigorous theoretical foundations and exhibits strong interpretability. Empirical evaluations across multimodal settings demonstrate its robustness and reliability, positioning it as a novel paradigm for dataset construction, augmentation, and evaluation.

dataset diversitydiversity metricsgeometric structure

HUMAP: Hierarchical Uniform Manifold Approximation and Projection

Jun 14, 2021
WE
Wilson E. Marc'ilio-Jr
🏛️ São Paulo State University (UNESP) | Nomic AI | Eindhoven University of Technology | Linnaeus University

Existing hierarchical dimensionality reduction methods struggle to simultaneously preserve both local and global structures across multiple granularity levels while maintaining users’ mental map continuity. To address this, we propose HUMAP—the first hierarchical dimensionality reduction framework built upon UMAP. HUMAP achieves dual optimization of structural fidelity and mental map consistency through three core techniques: cross-scale neighborhood graph propagation, multi-granularity similarity modeling, and incremental embedding alignment. Compared to state-of-the-art methods, HUMAP significantly improves hierarchical structure preservation across multiple benchmark datasets. Furthermore, it demonstrates strong interpretability and interactive utility in real-world data labeling tasks. By unifying geometric structure preservation with cognitive consistency, HUMAP establishes a novel paradigm for multi-granularity visual analytics, enabling scalable, intuitive, and semantically grounded exploration of high-dimensional hierarchical data.

Develops hierarchical dimensionality reduction for multi-granularity data analysisImproves mental map consistency compared to existing hierarchical approachesPreserves both local and global structures during hierarchical exploration

Latest Papers

What's happening recently
View more

From Topology to Retrieval: Decoding Embedding Spaces with Unified Signatures

Nov 27, 2025
FR
Florian Rottach
🏛️ University of Tübingen | The University of Texas at Austin | Fribourg University

This work addresses the weak interpretability of text embedding spaces and their limited structural representation. We propose the Unified Topological Signature (UTS) framework—the first systematic approach to jointly model the topological and geometric structure of embedding spaces. UTS integrates multi-dimensional features, including persistent homology, curvature estimation, and local density, overcoming the redundancy and low discriminability of conventional metrics. By applying clustering analysis and correlation modeling, UTS decodes the mapping between spatial organization and downstream retrieval performance, establishing a quantitative relationship between topological features and document retrievability. Extensive evaluation across multiple state-of-the-art embedding models and benchmark datasets demonstrates that UTS stably predicts inter-model performance differences and ranking effectiveness, exhibiting strong generalization capability and cross-model comparability.

Analyzing topological and geometric measures of text embedding spacesIntroducing a unified framework to characterize embedding spaces holisticallyLinking topological structure to retrieval performance and model properties

Predict Training Data Quality via Its Geometry in Metric Space

Oct 12, 2025
YB
Yang Ba
🏛️ Arizona State University

Conventional entropy-based diversity metrics fail to capture the intrinsic geometric and topological structure of high-dimensional training data, limiting their ability to assess data quality and predict model performance. Method: This work introduces persistent homology—a tool from topological data analysis—to quantify structural properties such as connected components, holes, and higher-order voids, thereby jointly characterizing data richness and redundancy through both topological and geometric lenses. Contribution/Results: Experiments demonstrate that the proposed topological metrics exhibit strong correlation with model generalization performance and serve as effective predictors of data quality. The framework enables principled data selection and efficient training, offering a novel paradigm for dataset curation. By leveraging topological signatures, it enhances training efficiency and robustness of AI systems without requiring model retraining or architectural modification. The approach is broadly applicable across domains where data geometry and topology critically influence learning dynamics.

Developing principled diversity measures beyond entropy-based metricsExploring topological features' impact on machine learning performanceQuantifying training data quality through geometric structure analysis

Traditional low-dimensional approaches struggle to capture the complex topological structure of neural network loss landscapes, limiting our understanding of optimization and generalization mechanisms. This work proposes Landscaper, an open-source tool that integrates Hessian-guided subspace sampling with topological data analysis (TDA) to enable geometric characterization of loss landscapes in arbitrary dimensions, revealing the hierarchy and connectivity of energy basins. The study introduces the Saddle-Minimum Average Distance (SMAD) metric, a novel indicator that, for the first time, leverages multidimensional topological analysis to detect landscape simplification during training. SMAD demonstrates sensitivity to training dynamics across diverse neural architectures and pretrained language models. Furthermore, in a molecular property prediction task, SMAD proves effective as a diagnostic indicator for out-of-distribution generalization.

generalizationhigh-dimensional geometryloss landscapes

This work proposes TopoPrune, a novel data pruning framework that addresses the instability of existing geometry-based methods under cross-architecture transfer or feature noise by leveraging topological inductive bias. TopoPrune uniquely integrates global low-dimensional manifold embedding with local differentiable persistent homology to enable dual-scale topological modeling: it first constructs a global manifold to capture the intrinsic data structure and then applies differentiable persistent homology to locally optimize samples and rank them by topological complexity. By exploiting the inherent stability of topological representations, TopoPrune maintains high accuracy even at aggressive pruning rates up to 90%, significantly outperforming current approaches while demonstrating exceptional robustness to noise and strong cross-architecture transferability.

cross-architecture transferdata pruningfeature noise

This work proposes a novel boundary extraction algorithm that addresses the challenge of recovering topologically consistent and watertight (hole-free) two-dimensional boundaries from multi-class implicit representations—a task where existing methods often fail. The method achieves, for the first time, topologically accurate and watertight boundary reconstruction under multi-class implicit settings, while incorporating a minimal detail constraint mechanism to control geometric approximation fidelity. Built upon implicit neural representations, the approach effectively preserves both boundary completeness and topological correctness. Experimental results on geological modeling datasets demonstrate that the algorithm accurately reconstructs complex topological structures, exhibiting strong adaptability and robustness across diverse scenarios.

boundary extractionimplicit representationsmulti-class

Hot Scholars

BR

Bastian Rieck

Professor, AIDOS Lab, University of Fribourg
Geometric Deep LearningTopological Data AnalysisTopological Deep Learning
ER

Eva Rotenberg

Associate Professor, DTU Compute, Denmark
AlgorithmsData StructuresGraph Algorithms
FF

Florent Foucaud

LIMOS, Université Clermont Auvergne, France
Graph theoryAlgorithmsComplexity
SH

Sariel Har-Peled

Professor of Computer Science, UIUC
Computational Geometry