Dataset Pruning from First Principles: A Label-Free Linear Programming Approach

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses geometric data pruning methods that rely on neighborhood similarity assumptions, which inherently introduce selection bias. Discarding this assumption, we reformulate unbiased subset selection from first principles as a variance minimization problem. Through a linear programming perspective, we construct high-dimensional polytopes and derive closed-form pairwise variance expressions, enabling an efficient vertex-walking algorithm for label-agnostic data pruning with strictly guaranteed statistical unbiasedness. Experiments across multiple benchmarks demonstrate that the proposed method outperforms uniform sampling and mainstream geometric approaches in accuracy, exhibiting particularly superior performance under small selection budgets while effectively reducing stochastic gradient descent (SGD) variance.
📝 Abstract
Dataset pruning reduces a large training set to a representative subset while preserving model performance. Existing geometry-based methods typically assume that nearby points in embedding space share similar properties. Rather than imposing this assumption, we derive geometric selection criteria by reformulating unbiased subset selection as a variance minimization problem. Unbiasedness ensures that unweighted subset averages recover full-dataset averages in expectation, including losses and gradients at fixed model parameters. Specifically, we characterize a family of unbiased subset selection algorithms as a high-dimensional polytope. In this context, minimizing the expected sampling variance is a linear objective. Differences in sampling variance, averaged over rigid motions, admit closed-form pairwise expressions. Because the polytope has high dimension, directly applying standard linear programming is impractical. We instead use these expressions to construct an efficient vertex walk that optimizes an approximation of the variance objective while preserving unbiasedness, yielding a method that requires neither labels nor model training during selection. Across CIFAR-10, MNIST, and CelebA benchmarks, our method matches or exceeds uniform sampling in mean test accuracy at every evaluated budget and outperforms competing geometric methods in several settings, particularly at small selection budgets. Beyond dataset pruning, the same framework reduces stochastic-gradient variance by increasing diversity within mini-batches while keeping the batch size unchanged.
Problem

Research questions and friction points this paper is trying to address.

Dataset Pruning
Subset Selection
Variance Minimization
Unbiasedness
Label-Free
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dataset Pruning
Label-Free Selection
Linear Programming
Variance Minimization
Vertex Walk
💼 Related Jobs
No related jobs found.
R
Rodrigo Schuller
Instituto de Matemática Pura e Aplicada (IMPA), Rio de Janeiro, Brazil
Francisco Ganacim
Francisco Ganacim
Instituto de Matemática Pura e Aplicada (IMPA), Rio de Janeiro, Brazil