CuBAS: Information Geometric Curvature-Based Adaptive Sampling for Supervised Classification

📅 2026-07-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of conventional sampling methods that overlook the geometric structure of data distributions, often yielding suboptimal training sets. The authors propose CuBAS, a novel framework that leverages local curvature from information geometry as a sampling criterion. Specifically, labeled data are modeled as a statistical manifold, and a q-state Potts Markov random field is employed to estimate local curvature via the ratio of first- and second-order Fisher information. Closed-form curvature scores are computed for each sample on a k-nearest neighbor graph, enabling adaptive selection of both low-curvature (homogeneous regions) and high-curvature (decision boundary-proximate) samples. CuBAS offers strong theoretical interpretability and practical efficacy, significantly outperforming random and uncertainty-based sampling across more than 60 benchmark datasets. It generalizes well across diverse classifiers and labeling budgets, with computational complexity scaling linearly with the number of graph edges.
📝 Abstract
The informativeness of a training set is as consequential as its size, yet most sampling strategies remain agnostic to the intrinsic geometry of the data distribution. We introduce CuBAS (Curvature-Based Adaptive Sampling), an information-geometric framework for adaptive data selection in supervised classification, grounded in the q-state Potts Markov random field (MRF) model. The central insight is that a labeled dataset can be viewed as a statistical manifold, on which local curvature, estimated via the ratio of second to first-order observed Fisher information, faithfully encodes the geometric complexity of the data distribution. We construct a k-nearest-neighbor graph over the labeled data and derive a closed-form curvature score at each vertex from the Potts sufficient statistics. This curvature signal partitions the graph into two complementary regimes: low-curvature regions, corresponding to smooth, homogeneous clusters, and high-curvature regions, concentrated around decision boundaries that are disproportionately informative for classification. By selecting nodes from both regimes, CuBAS constructs compact yet maximally informative training subsets. Empirical evaluation across more than 60 benchmark datasets demonstrates consistent and statistically significant improvements over random sampling and uncertainty-based baselines, across a wide range of labeling budgets and classifier architectures. CuBAS is computationally efficient (linear in the number of k-NN graph edges), theoretically grounded in the differential geometry of statistical manifolds, and interpretable in terms of the local shape operator of the data manifold.
Problem

Research questions and friction points this paper is trying to address.

adaptive sampling
supervised classification
information geometry
data manifold curvature
training set informativeness
Innovation

Methods, ideas, or system contributions that make the work stand out.

information geometry
curvature-based sampling
statistical manifold
adaptive data selection
Fisher information
🔎 Similar Papers
No similar papers found.
A
Alexandre L. M. Levada
Federal University of São Carlos