🤖 AI Summary
This work addresses the scalability challenge of pairwise loss functions in large-scale machine learning, whose computational complexity grows quadratically with data size. The authors propose a survey sampling–based estimator that directly samples pairs with carefully designed inclusion probabilities, leveraging auxiliary information to assign higher sampling weights to more informative pairs. This approach substantially reduces computational overhead while preserving optimization performance. Theoretical analysis establishes a precise trade-off bound between estimation accuracy and computational efficiency. Empirical results demonstrate that, in high-dimensional embedding tasks such as visual representation learning and graph learning, the method achieves performance comparable to full pairwise computation using only a small fraction of sampled pairs, thereby validating its effectiveness and scalability.
📝 Abstract
Many machine learning problems, including similarity learning, ranking, and clustering, rely on empirical pairwise loss functions whose quadratic computational cost quickly becomes prohibitive at scale. We demonstrate how a frugal approach that retains only a fraction of the available information on pairs can achieve estimation or optimization performance comparable to that obtained by using all pairs, by leveraging survey sampling techniques. A central finding, supported by both theory and experiments, is that such sampling plans must target pairs directly rather than individual observations. In particular, for pairwise losses between high-dimensional vectors such as embeddings in vision or graph learning, assigning higher inclusion probabilities to informative pairs using suitable auxiliary information yields performance close to full pairwise evaluation, providing a principled and theoretically grounded trade-off between accuracy and computational cost.