Doing well with less! On Sampling Techniques for Empirical Pairwise Loss Estimation/Minimization

📅 2026-06-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the scalability challenge of pairwise loss functions in large-scale machine learning, whose computational complexity grows quadratically with data size. The authors propose a survey sampling–based estimator that directly samples pairs with carefully designed inclusion probabilities, leveraging auxiliary information to assign higher sampling weights to more informative pairs. This approach substantially reduces computational overhead while preserving optimization performance. Theoretical analysis establishes a precise trade-off bound between estimation accuracy and computational efficiency. Empirical results demonstrate that, in high-dimensional embedding tasks such as visual representation learning and graph learning, the method achieves performance comparable to full pairwise computation using only a small fraction of sampled pairs, thereby validating its effectiveness and scalability.
📝 Abstract
Many machine learning problems, including similarity learning, ranking, and clustering, rely on empirical pairwise loss functions whose quadratic computational cost quickly becomes prohibitive at scale. We demonstrate how a frugal approach that retains only a fraction of the available information on pairs can achieve estimation or optimization performance comparable to that obtained by using all pairs, by leveraging survey sampling techniques. A central finding, supported by both theory and experiments, is that such sampling plans must target pairs directly rather than individual observations. In particular, for pairwise losses between high-dimensional vectors such as embeddings in vision or graph learning, assigning higher inclusion probabilities to informative pairs using suitable auxiliary information yields performance close to full pairwise evaluation, providing a principled and theoretically grounded trade-off between accuracy and computational cost.
Problem

Research questions and friction points this paper is trying to address.

pairwise loss
sampling
computational cost
empirical estimation
machine learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

pairwise loss
survey sampling
frugal learning
subsampling
computational efficiency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Louise Davy
IDS, LTCI, Télécom Paris, Palaiseau, France
S
Stephan Clémençon
IDS, LTCI, Télécom Paris, Palaiseau, France
C
Charlotte Laclau
IDS, LTCI, Télécom Paris, Palaiseau, France