Range Retrieval with Graph-Based Indices

📅 2025-02-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Range search—retrieving all points within distance ≤ r of a query point—in high-dimensional vector spaces remains a critical yet underexplored problem, with broad applications in duplicate detection, plagiarism checking, and face recognition. Method: This paper proposes an adaptive graph-based search framework built upon state-of-the-art graph indices (e.g., HNSW, NSG). It innovatively integrates distance distribution modeling, radius-adaptive selection, and dynamic resource scheduling to jointly optimize traversal strategies and termination criteria—effectively handling both empty-result and large-result scenarios. Contribution/Results: We introduce the first comprehensive range-search benchmark spanning multiple embedding scales, along with calibrated radius recommendations. On billion-scale datasets, our method achieves up to 100× higher throughput and 5–10× average speedup over baselines, demonstrating strong scalability and practical efficacy at the 100M-point scale.

Technology Category

Search and Optimization: Distributed SearchMachine Learning: Graph-based Machine LearningData Mining & Knowledge Management: Scalability, Parallel & Distributed Systems

Application Category

Graph Algorithms and Modeling for the Web: Querying, indexing, and retrieval in Web-related graphsSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesSemantics and Knowledge: Scalable techniques for the creation, curation, publication, maintenance, and consumption of large, Web-based, structured, reusable, knowledge graphs and ontologies
📝 Abstract
Retrieving points based on proximity in a high-dimensional vector space is a crucial step in information retrieval applications. The approximate nearest neighbor search (ANNS) problem, which identifies the $k$ nearest neighbors for a query (approximately, since exactly is hard), has been extensively studied in recent years. However, comparatively little attention has been paid to the related problem of finding all points within a given distance of a query, the range retrieval problem, despite its applications in areas such as duplicate detection, plagiarism checking, and facial recognition. In this paper, we present a set of algorithms for range retrieval on graph-based vector indices, which are known to achieve excellent performance on ANNS queries. Since a range query may have anywhere from no matching results to thousands of matching results in the database, we introduce a set of range retrieval algorithms based on modifications of the standard graph search that adapt to terminate quickly on queries in the former group, and to put more resources into finding results for the latter group. Due to the lack of existing benchmarks for range retrieval, we also undertake a comprehensive study of range characteristics of existing embedding datasets, and select a suitable range retrieval radius for eight existing datasets with up to 100 million points in addition to the one existing benchmark. We test our algorithms on these datasets, and find up to 100x improvement in query throughput over a naive baseline approach, with 5-10x improvement on average, and strong performance up to 100 million data points.
Problem

Research questions and friction points this paper is trying to address.

Range retrieval in high-dimensional spaces
Graph-based algorithms for efficient search
Adaptive techniques for varying result sizes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph-based vector indices
Adaptive range retrieval algorithms
Comprehensive dataset range study
🔎 Similar Papers
No similar papers found.
M
Magdalen Dobson Manohar
Carnegie Mellon University and Microsoft Azure, USA
T
Taekseung Kim
Carnegie Mellon University, USA
G
Guy E. Belloch
Carnegie Mellon University and Google, USA