Score
Designs, implements, and evaluates data structures, algorithms, and indexing schemes for fast approximate nearest-neighbor retrieval—covering k‑NN queries, nearest‑neighbor graphs, candidate generation, and search/pruning methods that balance recall, latency, and memory. Builds and analyzes distance-bounding and upper‑bounding techniques (including decoder‑informed bounds and structural lifts), efficient nearest‑neighbor indexing, and scalability strategies to handle large collections of vectors.
Managing and retrieving high-dimensional vector data poses significant challenges, particularly as traditional databases fail to meet performance requirements and the need for tight integration with large language models (LLMs) intensifies. Method: This paper systematically surveys four major approximate nearest neighbor search (ANNS) paradigms—hashing, tree-based indexing, graph-based methods (e.g., HNSW), and quantization (PQ/SQ)—and integrates hybrid optimization strategies. Contribution/Results: It introduces, for the first time, a “Four-Dimensional Methodology” framework tailored for industrial deployment of vector databases, analyzing trade-offs among accuracy, latency, memory footprint, and scalability. The work constructs a structured knowledge graph covering 200+ ANNS algorithms and proposes a novel paradigm for deep synergy between vector databases and LLMs. Collectively, these contributions provide both theoretical foundations and practical guidelines for system selection, architectural design, and development of AI-native database systems.
Sparse neighborhood graphs (SNGs) enable efficient approximate nearest neighbor search (ANNS) but lack rigorous theoretical foundations; existing truncation strategies are heuristic and often yield suboptimal performance. Method: This work introduces the first martingale-based analysis of the graph construction process, establishing tight theoretical guarantees: an $O(n^{2/3+varepsilon})$ upper bound on vertex degree and an $O(log n)$ bound on search path length. Leveraging these bounds, we propose a principled, theory-driven method for optimizing truncation parameters. Contribution/Results: Our approach significantly improves graph structural design and index construction. On billion-scale datasets, it achieves comparable or lower query latency while preserving Recall@10, and accelerates index building by 2–9×—thereby bridging the long-standing gap between theory and practice in graph-based ANNS.
This work addresses the challenge of efficient approximate nearest neighbor search and maximum inner product search (MIPS) in high-dimensional embeddings from large language models. The authors propose a novel approach that transforms asymmetric MIPS into Euclidean nearest neighbor search via dimension augmentation, combined with Equi-Voronoi Polytopes (EVP) quantization and a Fast Linear Assignment Sorting (FLAS) one-dimensional pre-sorting mechanism. This integration substantially accelerates k-nearest neighbor graph (kNNG) construction and query processing while enhancing memory access locality and cache efficiency. Evaluated in the SISAP 2026 challenge, the method achieves low-latency, high-recall MIPS performance, significantly outperforming existing baselines.
Range search—retrieving all points within distance ≤ r of a query point—in high-dimensional vector spaces remains a critical yet underexplored problem, with broad applications in duplicate detection, plagiarism checking, and face recognition. Method: This paper proposes an adaptive graph-based search framework built upon state-of-the-art graph indices (e.g., HNSW, NSG). It innovatively integrates distance distribution modeling, radius-adaptive selection, and dynamic resource scheduling to jointly optimize traversal strategies and termination criteria—effectively handling both empty-result and large-result scenarios. Contribution/Results: We introduce the first comprehensive range-search benchmark spanning multiple embedding scales, along with calibrated radius recommendations. On billion-scale datasets, our method achieves up to 100× higher throughput and 5–10× average speedup over baselines, demonstrating strong scalability and practical efficacy at the 100M-point scale.
The field of Fast Approximate Nearest Neighbor Search (FANNS) lacks a systematic survey addressing vector-scalar hybrid data, suffering from inconsistent problem formulations, absence of a unified algorithm taxonomy, and insufficient analysis of query difficulty. Method: We formally define hybrid datasets and hybrid queries, propose a fine-grained algorithm taxonomy centered on pruning mechanisms, and develop a distribution-sensitive query difficulty model. We further design a standardized evaluation framework and an open-source toolchain (Python/PyTorch) supporting hybrid dataset construction, quantitative difficulty assessment, and fair algorithm comparison. Contribution/Results: This work delivers the first structured, comprehensive survey of FANNS for hybrid data—filling a critical research gap. It establishes foundational theoretical principles and practical tools, enabling rigorous analysis and reproducible advancement in hybrid-data nearest neighbor search.
This work addresses the challenge of achieving both high scalability and efficiency in approximate nearest neighbor search (ANNS) over sparse vectors on conventional CPU architectures. To this end, we propose SpANNS—the first near-memory computing architecture tailored for sparse ANNS—built upon the CXL Type-2 platform. SpANNS integrates a hybrid inverted index, query parsing, clustering-based filtering, and a compute-enabled DIMM co-processing mechanism to perform index traversal and distance computation efficiently near the data. Evaluated against the state-of-the-art CPU baseline, SpANNS achieves a speedup of 15.2× to 21.6×, substantially enhancing both performance and scalability for sparse vector retrieval.
This work addresses the challenge of efficiently supporting both vector similarity search and arbitrary attribute filtering in high-dimensional approximate nearest neighbor retrieval. The authors propose a lightweight graph-based indexing algorithm that seamlessly integrates attribute filtering into the graph traversal process, overcoming the efficiency bottlenecks of existing methods when handling unseen query vectors combined with complex attribute constraints. Experimental results on multiple real-world datasets demonstrate that the proposed approach significantly outperforms state-of-the-art techniques, achieving substantially faster query latency while maintaining high recall. The method thus offers a compelling balance among flexibility, efficiency, and scalability for hybrid vector-and-attribute search scenarios.
This work addresses the challenge of efficiently merging graph indices in distributed systems and real-time vector databases, a problem previously lacking systematic investigation. To this end, the authors propose FGIM, a general and efficient three-stage framework for graph index merging. FGIM first converts input navigable graphs (e.g., HNSW) into k-nearest neighbor graphs (k-NNGs), then enhances neighbor quality and graph connectivity through cross-query candidate extraction and k-NNG refinement, and finally reconstructs a high-quality navigable graph. Extensive experiments on six real-world datasets demonstrate that FGIM achieves up to 3.5× speedup over incremental HNSW construction and averages 7.9× acceleration compared to non-incremental baselines, while maintaining comparable or superior retrieval accuracy.
This work addresses the significant performance degradation of existing quantization-based approximate nearest neighbor (ANN) methods under large-k queries, primarily caused by inefficient top-k result collection and costly re-ranking. To overcome these limitations, the authors propose a Bucket-based Collector (BBC), which organizes candidate vectors into distance-based buckets to reduce both candidate maintenance overhead and final sorting costs. Additionally, they introduce two efficient re-ranking algorithms tailored to different quantization schemes, effectively minimizing the number of items requiring re-ranking and mitigating cache misses. Experimental results demonstrate that, at a recall@k of 0.95, BBC accelerates state-of-the-art quantization-based ANN methods by up to 3.8×.
This work addresses the limitation of existing vector retrieval systems, which support only simple numerical or categorical constraints and struggle with complex graph-structured filtering requirements. To bridge this gap, the authors propose DLH, the first method to integrate graph-range filtering into approximate nearest neighbor (ANN) search. DLH constructs distance-aware label sets and compresses them into Bloom filters to enable efficient graph-range filtering. Furthermore, it introduces a query-node hash index caching mechanism (DLH-M) that exploits query locality to accelerate retrieval. Experimental results demonstrate that DLH achieves up to a 70.3% increase in throughput across multiple datasets while maintaining recall above 98.5%, with only modest additional storage overhead.
This study addresses the lack of comprehensive, reproducible benchmarks for vector databases across retrieval quality, latency, throughput, and resource consumption. For the first time, it jointly evaluates seven systems—FAISS, Qdrant, Milvus, Weaviate, Chroma, pgvector, and LanceDB—across six datasets (including SIFT, GIST, MS MARCO, and GloVe) comprising over four million vectors, measuring fifteen performance and resource metrics. The evaluation reveals that FAISS achieves the highest single-node throughput (866 QPS), Weaviate delivers out-of-the-box recall exceeding 99%, Qdrant exhibits the lowest median latency (4.55 ms), and LanceDB constructs indexes fastest albeit with slightly lower retrieval quality. The authors open-source the complete benchmarking framework and propose practical guidelines for system selection.