curvature-based adaptive sampling

Design and build sampling algorithms that compute information‑geometric or graph‑based curvature measures on a dataset (e.g., local manifold curvature from k‑NN graphs or learned representations) to score, partition, and select representative instances for labeling or training. The output is compact, regime‑aware subsets or ranked examples that improve label efficiency and the informativeness of training data.

curvature-basedadaptivesampling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitation of conventional sampling methods that overlook the geometric structure of data distributions, often yielding suboptimal training sets. The authors propose CuBAS, a novel framework that leverages local curvature from information geometry as a sampling criterion. Specifically, labeled data are modeled as a statistical manifold, and a q-state Potts Markov random field is employed to estimate local curvature via the ratio of first- and second-order Fisher information. Closed-form curvature scores are computed for each sample on a k-nearest neighbor graph, enabling adaptive selection of both low-curvature (homogeneous regions) and high-curvature (decision boundary-proximate) samples. CuBAS offers strong theoretical interpretability and practical efficacy, significantly outperforming random and uncertainty-based sampling across more than 60 benchmark datasets. It generalizes well across diverse classifiers and labeling budgets, with computational complexity scaling linearly with the number of graph edges.

adaptive samplingdata manifold curvatureinformation geometry

Curvature as a tool for evaluating dimensionality reduction and estimating intrinsic dimension

Sep 16, 2025
CB
Charlotte Beylier
🏛️ Center for Scalable Data Analytics and Artificial Intelligence (ScaDS,AI) Dresden/Leipzig | Max Planck Institute for Mathematics in the Sciences | Max Planck Institute for Human Cognitive and Brain Sciences

This paper addresses the challenges of evaluating dimensionality reduction (DR) effectiveness and estimating intrinsic data dimensionality. We propose a geometric profiling method based on sectional curvature in discrete metric spaces, which characterizes large-scale data geometry via metric relationships among point triplets. For the first time, this approach systematically introduces differential-geometric curvature into quantitative DR quality assessment and intrinsic dimension estimation—without requiring embedded coordinates or manifold assumptions, thus ensuring both theoretical rigor and computational feasibility. Experiments across diverse synthetic and real-world datasets demonstrate that our method robustly discriminates DR algorithm performance, achieves significantly lower intrinsic dimension estimation error than state-of-the-art methods (e.g., MDS- and PCA-based estimators), and successfully uncovers latent negative curvature in empirical networks—including social and biological networks.

Estimating intrinsic dimensionality of datasets using curvature analysisEvaluating dimensionality reduction effectiveness via curvature profilesExploring large-scale geometry of networks with curvature-based methods

A roadmap for curvature-based geometric data analysis and learning

Oct 26, 2025
YY
Yasharth Yadav
🏛️ Nanyang Technological University

Discrete curvature modeling lacks a unified theoretical foundation and practical framework for geometric data analysis and learning. Method: This work establishes the first comprehensive, end-to-end technical roadmap for discrete curvature—unifying Riemannian and metric geometric perspectives—and introduces a curvature-driven multi-structural analysis framework compatible with graphs, simplicial complexes, cubical complexes, and point clouds. It integrates discrete differential geometry, topological data analysis (TDA), and geometric deep learning to formulate a standardized evaluation protocol. Contribution/Results: We present (i) the first holistic technical roadmap dedicated to discrete curvature; (ii) a unified theoretical analysis framework grounded in geometric principles; and (iii) an open-source, reproducible benchmark for comparative evaluation of multiple discrete curvature models. Empirical results demonstrate significant improvements in modeling intrinsic geometric structures of complex data and enhancing generalization across downstream tasks.

Providing computational algorithms for curvature-driven data processingReviewing discrete curvature models for geometric data analysisSurveying curvature applications in supervised and unsupervised learning

Recovering Manifold Structure Using Ollivier-Ricci Curvature

Oct 02, 2024
TL
Tristan Luca Saidi
🏛️ Columbia University

To address the problem of spurious inter-manifold edges in nearest-neighbor graphs—induced by noisy data—that distort low-dimensional manifold structure, this paper proposes ORC-ManL, a novel manifold learning algorithm. Methodologically, it is the first to incorporate Ollivier–Ricci curvature (ORC) into manifold learning, establishing a theoretical link between negative ORC and spurious edges; it further integrates local metric distortion estimation to enable geometry-driven, provably convergent graph pruning. The framework unifies ORC computation, topological optimization, and persistent homology evaluation. Empirically, ORC-ManL significantly outperforms existing pruning methods across manifold learning, intrinsic dimension estimation, and single-cell RNA-seq clustering tasks, yielding 15–32% improvements in downstream accuracy. Crucially, its theoretical convergence guarantee is empirically validated.

Enhance geometric data analysis tasks performanceImprove manifold learning with Ollivier-Ricci curvaturePrune spurious edges from nearest neighbor graphs

Statistical exploration of the Manifold Hypothesis

Aug 24, 2022
NW
N. Whiteley
🏛️ University of Bristol | University of Edinburgh

This work addresses the lack of a universal statistical interpretation for the manifold hypothesis—that high-dimensional data approximately reside on low-dimensional manifolds. We propose the Latent Metric Model (LMM), a generative framework grounded in fundamental statistical concepts: latent variables, variable dependence, and stationarity—providing the first unified statistical justification for the manifold assumption. Methodologically, LMM integrates neighborhood graph construction, spectral analysis, and an interpretable inference framework to enable unsupervised manifold discovery and geometric structure recovery under weak priors. Experiments demonstrate that complex manifold geometries naturally emerge from minimal statistical mechanisms; LMM significantly reduces reliance on hand-crafted priors on both synthetic and real-world datasets, while enabling interpretable reconstruction of manifold dimensionality, curvature, and coordinate systems.

Develops methods to discover and interpret high-dimensional data geometryExplores why high-dimensional data concentrates near low-dimensional manifoldsProposes Latent Metric Model to explain manifold structure emergence

Latest Papers

What's happening recently
View more

该研究提出了一种基于电阻曲率引导的子图采样框架ERC-LG,以解决大规模图神经网络训练成本高且采样标准忽略边几何角色的问题。

geometric roles of edgeslarge-scale graph neural networkssubgraph sampling

This work addresses geometric data pruning methods that rely on neighborhood similarity assumptions, which inherently introduce selection bias. Discarding this assumption, we reformulate unbiased subset selection from first principles as a variance minimization problem. Through a linear programming perspective, we construct high-dimensional polytopes and derive closed-form pairwise variance expressions, enabling an efficient vertex-walking algorithm for label-agnostic data pruning with strictly guaranteed statistical unbiasedness. Experiments across multiple benchmarks demonstrate that the proposed method outperforms uniform sampling and mainstream geometric approaches in accuracy, exhibiting particularly superior performance under small selection budgets while effectively reducing stochastic gradient descent (SGD) variance.

Dataset PruningLabel-FreeSubset Selection

Current evaluations in relational learning rely on unified leaderboards that overlook intrinsic geometric differences among datasets, often leading to misleading judgments of model generalization. This work proposes a curvature-stratified, geometry-aware evaluation framework, categorizing 14 datasets into positively curved, negatively curved, and near-zero curvature groups, and systematically assessing the performance of 18 models—including GCNs, graph foundation models, and tabular methods—across these geometric regimes. Experiments reveal for the first time that model performance is highly dependent on data geometry: rankings remain stable within each curvature regime but shift significantly across regimes. Notably, in certain settings, curvature-aligned GNNs even outperform graph foundation models. These findings challenge the universality assumption underlying conventional aggregate metrics and establish a more fine-grained, structure-aware paradigm for evaluating relational learning methods.

curvatureevaluation biasintrinsic geometry

This work addresses the poor generalization of machine learning models on variable-sized inputs—such as point clouds, sequences, and graphs—and the high computational cost of evaluating large inputs. To this end, the authors propose a unified framework based on randomized sampling that explicitly links sampling strategies to the symmetry properties of input structures. By integrating function continuity theory, they establish the first generalization and compression theory applicable across diverse structured data types, including sequences, graphs, and tensors, and derive explicit bounds on generalization error and sketching rates. The framework encompasses generalized forms of sampling with replacement, random binning, and species sampling, and is successfully applied to moment polynomials, graph homomorphism densities, permutation-invariant Transformers, and graph neural networks, enabling efficient cross-scale approximation with strong generalization guarantees.

generalizationinput sizemodel evaluation

Hot Scholars

TZ

Tianyi Zhang

Assistant Professor of Computer Science, Purdue University
Software EngineeringHuman-Computer InteractionLarge Language Models
CS

Christian S. Jensen

Aalborg University, Department of Computer Science
Data managementanalyticsindexingquery processing
ZJ

Zhi Jin

Sun Yat-Sen University, Associate Professor
XC

Xiaowen Chu

IEEE Fellow, Professor, Data Science and Analytics, HKUST(GZ)
GPU ComputingMachine Learning SystemsParallel and Distributed ComputingWireless Networks