Score
Designs and evaluates methods that estimate how well a pre-trained model or representation will transfer to a new task by analyzing the topological structure of data and feature spaces; this includes building non-parametric, topology-aware metrics and aligning sparse topological summaries (e.g., 1-skeleton graphs or other manifold descriptors) with label information to rank or predict transferability without full fine‑tuning.
Existing methods for evaluating transferability in 3D medical image segmentation rely on time-consuming fine-tuning, which struggles to meet the stringent demands for boundary precision and anatomical consistency. This work proposes the first fine-tuning-free, topology-driven framework that aligns sparse features with semantic labels via minimum spanning trees (MSTs). It assesses transferability at dual scales—local boundary token separability (LBTC) and global representation topological divergence (GRTD)—and incorporates a task-adaptive gated fusion mechanism. Theoretically, we prove that the MST leakage rate constitutes a finite-sample lower bound of the Bayes error and reveal that randomly initialized decoders stabilize topological alignment. Evaluated on a large-scale benchmark encompassing 114,000 3D medical images, our method achieves state-of-the-art performance, improving the weighted Kendall metric by 0.36 on average and accelerating evaluation by 56×.
Conventional entropy-based diversity metrics fail to capture the intrinsic geometric and topological structure of high-dimensional training data, limiting their ability to assess data quality and predict model performance. Method: This work introduces persistent homology—a tool from topological data analysis—to quantify structural properties such as connected components, holes, and higher-order voids, thereby jointly characterizing data richness and redundancy through both topological and geometric lenses. Contribution/Results: Experiments demonstrate that the proposed topological metrics exhibit strong correlation with model generalization performance and serve as effective predictors of data quality. The framework enables principled data selection and efficient training, offering a novel paradigm for dataset curation. By leveraging topological signatures, it enhances training efficiency and robustness of AI systems without requiring model retraining or architectural modification. The approach is broadly applicable across domains where data geometry and topology critically influence learning dynamics.
This work addresses the challenge of topological representation for large-scale geometric data, where existing methods—such as Mapper and geometric simplicial complexes—suffer from deteriorating performance at scale, heavy reliance on manual hyperparameter tuning, and limited expressivity in capturing high-dimensional topological features. We propose the first differentiable optimization framework that learns topology-preserving covers end-to-end, eliminating dependence on predefined geometric structures or hand-crafted design. Our core innovation is modeling the cover as a differentiable neural network and introducing topologically grounded loss functions—e.g., nerve consistency—to enable gradient-based parameterization and optimization of covering subsets. The resulting simplicial complexes are more compact yet exhibit superior preservation of high-dimensional topology. Extensive evaluation demonstrates significant improvements over traditional methods on large datasets, with enhanced scalability, robustness to data scale, and fidelity to underlying topological structure.
This work addresses the high cost of selecting self-supervised pre-trained models for medical image segmentation by proposing a fine-tuning-free, topology-driven transferability evaluation framework. It introduces manifold topology into the assessment of medical foundation models, quantifying both global structural isomorphism—via minimum spanning trees and manifold separability—and local boundary topological consistency between features and label manifolds. A task-adaptive fusion mechanism is further designed to rank candidate models effectively. Evaluated on the OpenMind benchmark, the proposed method achieves approximately a 31% improvement in the weighted Kendall metric over existing approaches, significantly enhancing both the efficiency and accuracy of model selection.
In neuroimaging research, the diversity of rfMRI brain representations—e.g., parcellation schemes and feature types—induces inconsistent individual variability estimates, severely undermining result reproducibility and cross-study integration. Method: We propose a topological similarity assessment framework based on persistent homology, specifically designed for small-sample, high-dimensional neuroimaging data. It integrates Vietoris–Rips complexes, topological bootstrapping, and hierarchical clustering, and introduces a novel prevalence-weighted Wasserstein distance to enable unbiased topological comparison across samples and heterogeneous metric spaces. Contribution/Results: We demonstrate that low-persistence but high-prevalence homological generators encode interpretable biological signals. Validation on large-scale cohorts reveals that environmental dimensions of representations—not decomposition rank—predominantly govern topological feature count and stability. The framework enables robust, unbiased clustering across diverse representations, advancing reproducible, integrative neuroimaging analysis.
This work addresses the lack of theoretical foundations for substructure transferability in graph data by bridging transferable substructures with the intrinsic geometry of graph representation spaces from a functional behavior perspective. It proposes the first Riemannian geometry–based framework for learning intrinsic graph geometry, innovatively introducing neural vector bundles and local coordinate charts to construct the GAUGE pretraining architecture. A Dirichlet loss function is designed to enable explicit modeling of intrinsic graph geometry and quantification of transfer difficulty. The method demonstrates significant performance gains over existing models on zero-shot link prediction and graph isomorphism tasks, validating its expressive power and cross-task transferability.
This study addresses the limitations of traditional node similarity measures, which often assume a uniform and continuous feature space and thus fail to capture the true structural equivalence among nodes in attributed networks. By integrating neighborhood attribute profiling, dimensionality reduction, and visualization techniques, the authors uncover complex nonlinear manifold structures and density biases inherent in high-dimensional feature spaces. Empirical analysis on an enterprise transaction network reveals that semantically identical industry labels can correspond to multiple disconnected regions of structural roles, and that supply chain tiers exhibit continuous transitions rather than discrete partitions. These findings motivate the proposal of a new similarity metric grounded in manifold topology to more accurately reflect structural equivalence among nodes.