Score
Designs, builds, and evaluates systems that store, index, and retrieve high‑dimensional numeric vectors to provide low‑latency, high‑recall similarity search and nearest‑neighbor retrieval over large collections. Work includes selecting and implementing vector indexes and ANN/quantization algorithms, data ingestion/update and persistence strategies, sharding/replication and resource management, hybrid scalar+vector query support, integration with embedding pipelines and application stacks, and tuning latency/throughput/accuracy tradeoffs and operational monitoring.
Managing and retrieving high-dimensional vector data poses significant challenges, particularly as traditional databases fail to meet performance requirements and the need for tight integration with large language models (LLMs) intensifies. Method: This paper systematically surveys four major approximate nearest neighbor search (ANNS) paradigms—hashing, tree-based indexing, graph-based methods (e.g., HNSW), and quantization (PQ/SQ)—and integrates hybrid optimization strategies. Contribution/Results: It introduces, for the first time, a “Four-Dimensional Methodology” framework tailored for industrial deployment of vector databases, analyzing trade-offs among accuracy, latency, memory footprint, and scalability. The work constructs a structured knowledge graph covering 200+ ANNS algorithms and proposes a novel paradigm for deep synergy between vector databases and LLMs. Collectively, these contributions provide both theoretical foundations and practical guidelines for system selection, architectural design, and development of AI-native database systems.
This work addresses the lack of systematic design principles for neural retrieval systems that balance efficiency and effectiveness. It proposes the first vertically layered four-tier framework—spanning representation, granularity, orchestration, and robustness—to structurally characterize key design decisions at each layer and their interdependencies. By integrating Bi- and Cross-encoder architectures, atomic and hierarchical chunking strategies, multi-stage re-ranking, agent-based decomposition, and domain generalization techniques, the study elucidates the mechanistic impact of each design choice on system performance. This approach effectively mitigates critical challenges such as information bottlenecks, semantic blind spots, and temporal drift, thereby offering a practical and actionable optimization pathway for building efficient and robust embedded retrieval systems.
This study addresses the lack of comprehensive, reproducible benchmarks for vector databases across retrieval quality, latency, throughput, and resource consumption. For the first time, it jointly evaluates seven systems—FAISS, Qdrant, Milvus, Weaviate, Chroma, pgvector, and LanceDB—across six datasets (including SIFT, GIST, MS MARCO, and GloVe) comprising over four million vectors, measuring fifteen performance and resource metrics. The evaluation reveals that FAISS achieves the highest single-node throughput (866 QPS), Weaviate delivers out-of-the-box recall exceeding 99%, Qdrant exhibits the lowest median latency (4.55 ms), and LanceDB constructs indexes fastest albeit with slightly lower retrieval quality. The authors open-source the complete benchmarking framework and propose practical guidelines for system selection.
To address high query latency and low recall in approximate nearest neighbor (ANN) search under dynamic skew workloads, this paper proposes AdaptANN—a self-adaptive indexing framework designed for high-dimensional vector streams with frequent updates and queries. Its core contributions are threefold: (1) a novel hierarchical adaptive partitioning mechanism enabling localized re-indexing in response to evolving data distributions; (2) a NUMA-aware parallel query engine coupled with an online access-frequency-driven cost-recall joint model for real-time query parameter optimization; and (3) a lightweight recall estimator guaranteeing target recall. Evaluated on the dynamic Wikipedia vector benchmark, AdaptANN achieves 1.5–22× lower query latency and 6–83× lower index update latency compared to SVS, DiskANN, HNSW, and SCANN, significantly improving both efficiency and accuracy in dynamic settings.
Existing graph-based indexes for vector search overlook the spatial and temporal locality inherent in query streams, leading to redundant traversals and suboptimal efficiency. This work proposes CatapultDB, the first approach to dynamically model query locality within ANN graph indexes and inject lightweight “catapult” shortcut edges that connect high-frequency query regions to relevant target nodes. By doing so, CatapultDB selects superior starting points for queries without altering the underlying graph structure or search algorithm. The method is transparently compatible with filtered search, dynamic insertions, and disk-resident indexes. Experimental results on four skewed workloads show that CatapultDB achieves up to 2.51× higher throughput than DiskANN while matching or exceeding its recall, delivering efficiency comparable to LSH—without requiring index reconstruction or sacrificing functionality—and gracefully adapts to shifting query patterns.
This work addresses a critical limitation in existing graph-based disk indexing systems for large-scale high-dimensional vector similarity search: their performance is constrained by overlooking computational overhead, as the true bottleneck in high-dimensional settings lies in computation rather than I/O. The study is the first to reveal the intrinsic nature of this computational bottleneck and proposes a novel computation-optimized disk data layout that fully exploits modern CPU SIMD instructions. The approach integrates degree-based node caching, cluster-driven entry point selection, and an early scheduling strategy. Extensive experiments demonstrate that the proposed method significantly outperforms state-of-the-art disk-based graph index systems across multiple large-scale high-dimensional datasets, achieving performance comparable to—or even surpassing—that of in-memory indexing schemes, thereby transcending the traditional I/O-centric design paradigm.
This work addresses the challenge of achieving both high scalability and efficiency in approximate nearest neighbor search (ANNS) over sparse vectors on conventional CPU architectures. To this end, we propose SpANNS—the first near-memory computing architecture tailored for sparse ANNS—built upon the CXL Type-2 platform. SpANNS integrates a hybrid inverted index, query parsing, clustering-based filtering, and a compute-enabled DIMM co-processing mechanism to perform index traversal and distance computation efficiently near the data. Evaluated against the state-of-the-art CPU baseline, SpANNS achieves a speedup of 15.2× to 21.6×, substantially enhancing both performance and scalability for sparse vector retrieval.
This work addresses the performance, latency, scalability, and cost challenges posed by the exponential growth of vector data by systematically tracing the evolution of vector search technologies through the lens of storage architecture. It proposes a cloud-native three-tier storage framework—comprising memory, SSD, and object storage—that integrates in-memory indexing methods such as IVF, hashing, quantization, and graph-based indices with heterogeneous storage techniques including block-level data layout, I/O optimization, and efficient index updates. Furthermore, it introduces a cloud-native data tiering strategy to orchestrate data placement across storage layers. The resulting framework provides a comprehensive theoretical foundation and practical guidance for building trillion-scale vector retrieval systems that achieve high throughput, low latency, and cost efficiency, while also outlining key directions for future research.
This study addresses the lack of systematic evaluation of hybrid search mechanisms that combine semantic retrieval with metadata filtering in existing vector databases. We propose a novel relevance metric, Global-Local Selectivity (GLS), construct MoReVec—the first benchmark dataset supporting filtered retrieval—and extend ANN-Benchmarks to enable unified evaluation of hybrid search performance. Through comprehensive experiments integrating diverse filtering strategies into FAISS, Milvus, and pgvector with IVFFlat and HNSW indexes, we demonstrate that engine-level algorithmic integration critically governs performance: Milvus achieves more stable recall via hybrid execution, pgvector’s optimizer often selects suboptimal query plans, and IVFFlat outperforms HNSW under low-selectivity queries. Our findings culminate in practical configuration guidelines that offer both theoretical insights and actionable recommendations for efficient hybrid search deployment.
This paper addresses performance bottlenecks of cloud-native vector search over remote storage by systematically comparing cluster-based and graph-based indexes under high-concurrency, high-recall, and high-dimensional workloads. Through a unified benchmarking framework, we reveal the graph index’s substantial advantages in latency-sensitive and cache-constrained scenarios, and identify the root cause of local tuning parameter failure in cloud environments: a fundamental mismatch between index access granularity and cloud storage caching mechanisms. To resolve this, we propose a cloud-aware graph index parameter redesign method and an index-cache co-optimization strategy that dynamically adapts query granularity and data-fetch patterns to available cache capacity. Experimental evaluation on representative cloud configurations demonstrates an average 37% reduction in query latency and a 5.2% improvement in recall. This work is the first to rigorously characterize the cloud-native applicability boundaries of these two dominant indexing paradigms.
This work addresses the inefficiency of existing vector databases under high-recall and low-latency requirements, where memory and I/O bottlenecks hinder effective query-level caching. The authors propose a backend-agnostic, semantic-aware caching system that introduces, for the first time, a vector cache layer capable of handling semantically redundant queries. By dynamically adjusting region-wise distance thresholds through online learning, the system achieves high cache hit rates within a bounded memory footprint of just a few megabytes. Deployed as a plug-in component, it substantially reduces end-to-end query latency by 40–1000×, with cache-hit responses under 1 millisecond, while preserving recall performance comparable to that of the underlying approximate nearest neighbor (ANN) search backend.
This work addresses the high latency and service disruption caused by frequent index rebuilds in existing approximate nearest neighbor (ANN) methods under dynamic vector database updates. To overcome these limitations, the authors propose ACRONYM—a co-designed algorithm-hardware platform that leverages a data-distribution-agnostic XOR-and-Accumulate (XAC) systolic array encoder and Hamming-distance-based search, integrated with content-addressable memory (CAM) to enable in-memory parallel computation. A two-stage coarse-to-fine retrieval architecture circumvents CAM dimensionality constraints, allowing continuous, interruption-free updates. Evaluated on million-scale dynamic datasets, ACRONYM achieves over 90% recall, 8 million queries per second throughput, only 32 MB memory footprint, and 2.56 μJ per query energy efficiency—outperforming CPU-based HNSW by 400× and GPU-based FAISS-IVF by 80× in speed.