Score
Extracting and using multi-scale topological or geometric summaries to represent feature evolution, capture forms of ill-posedness, and serve as surrogates for unobserved concept changes in models.
This work addresses the fundamental challenge that 2D vision models cannot be directly applied to irregular and sparse 3D data such as point clouds and meshes. To this end, it proposes the first unified taxonomy encompassing data representations, architectural designs, and hybrid strategies. Existing approaches are systematically categorized into three paradigms: projection-based data-centric methods, architecture-centric techniques leveraging native 3D structures, and hybrid approaches combining both. The study provides a thorough analysis of the trade-offs among computational complexity, reliance on pretraining, and preservation of geometric inductive biases. By integrating pathways including transfer from 2D CNNs or Vision Transformers, native 3D network design, multimodal fusion, and self-supervised learning, this work offers a systematic roadmap for advancing 3D foundation models, geometric self-supervised learning, and multimodal representation learning.
Traditional machine learning struggles to effectively model shape data with nonlinear geometric structures and their intrinsic variability. This work proposes a unified analytical framework that systematically integrates differential geometry, manifold statistics, and geometric deep learning to address the challenges posed by complex, unaligned shapes exhibiting nonlinear variation. The framework encompasses key components including shape representation, geodesic metrics, parametrization, and statistical inference. It has been successfully applied to multiscale biological geometric data—such as cellular morphologies and primate dental evolution—revealing structural patterns and evolutionary trajectories underlying shape variation. This approach establishes both a theoretical foundation and a practical paradigm for geometry-aware learning in shape analysis.
This work addresses the challenge of topological representation for large-scale geometric data, where existing methods—such as Mapper and geometric simplicial complexes—suffer from deteriorating performance at scale, heavy reliance on manual hyperparameter tuning, and limited expressivity in capturing high-dimensional topological features. We propose the first differentiable optimization framework that learns topology-preserving covers end-to-end, eliminating dependence on predefined geometric structures or hand-crafted design. Our core innovation is modeling the cover as a differentiable neural network and introducing topologically grounded loss functions—e.g., nerve consistency—to enable gradient-based parameterization and optimization of covering subsets. The resulting simplicial complexes are more compact yet exhibit superior preservation of high-dimensional topology. Extensive evaluation demonstrates significant improvements over traditional methods on large datasets, with enhanced scalability, robustness to data scale, and fidelity to underlying topological structure.
Prior work designs geometric structures for specific features, lacking generalizability across model families and scales. Method: We propose Supervised Multidimensional Scaling (SMDS), the first method to automatically discover feature manifolds across diverse large language models (LLMs) and sizes, systematically characterizing the geometric representations of concepts in activation space. By integrating concept-labeled data into manifold learning, SMDS uncovers distinct geometric structures—circular, linear, and cluster-like—in temporal reasoning, corresponding respectively to periodicity, sequentiality, and categorical semantics. Results: We empirically demonstrate that these structures dynamically reconfigure within context and functionally support concretized reasoning. Experiments confirm their cross-model stability, interpretability, and strong correlation with downstream reasoning performance.
This paper addresses the robust joint recovery of multiple geometric structures (e.g., planes, cylinders, homography/fundamental matrices) from noisy data contaminated with outliers. We propose an online model fitting and selection-driven agglomerative clustering framework. Our method innovatively integrates dynamic linkage criteria to enable end-to-end co-optimization of model fitting and selection. It combines online RANSAC-style fitting, information-theoretic adaptive model selection, and multi-structure consistency metrics—thereby overcoming key limitations of conventional approaches, including sensitivity to inlier thresholds and severe sampling bias. Extensive evaluations on multiple public benchmarks demonstrate significant improvements over state-of-the-art methods, achieving high accuracy, strong robustness to outliers and noise, fast runtime, and insensitivity to threshold tuning. The source code is publicly available.
Existing diversity metrics for datasets predominantly rely on statistical distributions or entropy, often overlooking the intrinsic geometric structure of the data. This work introduces persistence landscapes (PLs)—a tool from topological data analysis—into diversity assessment, offering a geometric perspective to quantify structural diversity and establishing a direct link between geometric features and diversity. The proposed PLDiv metric is grounded in rigorous theoretical foundations and exhibits strong interpretability. Empirical evaluations across multimodal settings demonstrate its robustness and reliability, positioning it as a novel paradigm for dataset construction, augmentation, and evaluation.
Existing hierarchical dimensionality reduction methods struggle to simultaneously preserve both local and global structures across multiple granularity levels while maintaining users’ mental map continuity. To address this, we propose HUMAP—the first hierarchical dimensionality reduction framework built upon UMAP. HUMAP achieves dual optimization of structural fidelity and mental map consistency through three core techniques: cross-scale neighborhood graph propagation, multi-granularity similarity modeling, and incremental embedding alignment. Compared to state-of-the-art methods, HUMAP significantly improves hierarchical structure preservation across multiple benchmark datasets. Furthermore, it demonstrates strong interpretability and interactive utility in real-world data labeling tasks. By unifying geometric structure preservation with cognitive consistency, HUMAP establishes a novel paradigm for multi-granularity visual analytics, enabling scalable, intuitive, and semantically grounded exploration of high-dimensional hierarchical data.
This work addresses the weak interpretability of text embedding spaces and their limited structural representation. We propose the Unified Topological Signature (UTS) framework—the first systematic approach to jointly model the topological and geometric structure of embedding spaces. UTS integrates multi-dimensional features, including persistent homology, curvature estimation, and local density, overcoming the redundancy and low discriminability of conventional metrics. By applying clustering analysis and correlation modeling, UTS decodes the mapping between spatial organization and downstream retrieval performance, establishing a quantitative relationship between topological features and document retrievability. Extensive evaluation across multiple state-of-the-art embedding models and benchmark datasets demonstrates that UTS stably predicts inter-model performance differences and ranking effectiveness, exhibiting strong generalization capability and cross-model comparability.
Conventional entropy-based diversity metrics fail to capture the intrinsic geometric and topological structure of high-dimensional training data, limiting their ability to assess data quality and predict model performance. Method: This work introduces persistent homology—a tool from topological data analysis—to quantify structural properties such as connected components, holes, and higher-order voids, thereby jointly characterizing data richness and redundancy through both topological and geometric lenses. Contribution/Results: Experiments demonstrate that the proposed topological metrics exhibit strong correlation with model generalization performance and serve as effective predictors of data quality. The framework enables principled data selection and efficient training, offering a novel paradigm for dataset curation. By leveraging topological signatures, it enhances training efficiency and robustness of AI systems without requiring model retraining or architectural modification. The approach is broadly applicable across domains where data geometry and topology critically influence learning dynamics.
Traditional low-dimensional approaches struggle to capture the complex topological structure of neural network loss landscapes, limiting our understanding of optimization and generalization mechanisms. This work proposes Landscaper, an open-source tool that integrates Hessian-guided subspace sampling with topological data analysis (TDA) to enable geometric characterization of loss landscapes in arbitrary dimensions, revealing the hierarchy and connectivity of energy basins. The study introduces the Saddle-Minimum Average Distance (SMAD) metric, a novel indicator that, for the first time, leverages multidimensional topological analysis to detect landscape simplification during training. SMAD demonstrates sensitivity to training dynamics across diverse neural architectures and pretrained language models. Furthermore, in a molecular property prediction task, SMAD proves effective as a diagnostic indicator for out-of-distribution generalization.
This work proposes TopoPrune, a novel data pruning framework that addresses the instability of existing geometry-based methods under cross-architecture transfer or feature noise by leveraging topological inductive bias. TopoPrune uniquely integrates global low-dimensional manifold embedding with local differentiable persistent homology to enable dual-scale topological modeling: it first constructs a global manifold to capture the intrinsic data structure and then applies differentiable persistent homology to locally optimize samples and rank them by topological complexity. By exploiting the inherent stability of topological representations, TopoPrune maintains high accuracy even at aggressive pruning rates up to 90%, significantly outperforming current approaches while demonstrating exceptional robustness to noise and strong cross-architecture transferability.
This work proposes a novel boundary extraction algorithm that addresses the challenge of recovering topologically consistent and watertight (hole-free) two-dimensional boundaries from multi-class implicit representations—a task where existing methods often fail. The method achieves, for the first time, topologically accurate and watertight boundary reconstruction under multi-class implicit settings, while incorporating a minimal detail constraint mechanism to control geometric approximation fidelity. Built upon implicit neural representations, the approach effectively preserves both boundary completeness and topological correctness. Experimental results on geological modeling datasets demonstrate that the algorithm accurately reconstructs complex topological structures, exhibiting strong adaptability and robustness across diverse scenarios.