Score
Designs and implements index structures and complete indexing systems for efficient storage, lookup, and retrieval across modalities (e.g., full-text, hash-based, spatial, graph, and vector/similarity search), producing compact searchable representations and scalable I/O. Builds and analyzes index construction and update pipelines—including incremental updates, ANN/vector index construction, semantic and similarity indexing—and optimizes indexing algorithms and structures for lookup, retrieval latency, throughput, and resource trade-offs.
To address the prohibitively long indexing time and poor CPU-architecture compatibility of graph-based indexes in high-dimensional approximate nearest neighbor search (ANNS), this paper proposes Flash—a hardware-aware compact encoding strategy. Flash innovatively integrates vector quantization with key CPU architectural features, including SIMD parallelism and cache locality, enabling efficient distance computation while maintaining bounded quantization error. By jointly optimizing compact encoding, memory access patterns, and cache-friendly graph construction, Flash achieves 10.4×–22.9× speedup in index construction across eight real-world datasets ranging from 10 million to one billion vectors. Crucially, this acceleration comes without sacrificing retrieval accuracy or query latency—indeed, both are preserved or improved. Flash thus bridges the gap between algorithmic efficiency and modern hardware utilization in large-scale ANNS.
This work addresses the challenge of balancing recall accuracy and throughput efficiency under frequent vector updates in the era of large language models. The authors propose Yi, a novel graph-based vector indexing system that pioneers support for in-place updates. Guided by a “decompose-to-integrate” design philosophy, Yi introduces a vector-level update mechanism that overcomes the bottlenecks of traditional batch merging. It integrates three core components—a tasklet execution engine, an asynchronous buffer manager, and a vector file system—to enable efficient online updates and high-quality retrieval. Experiments on a dataset of 800 million vectors demonstrate that Yi achieves 1.75× higher update throughput, 1.8× greater concurrent search throughput, and reduces peak memory usage to 73% of baseline systems, delivering superior performance even with fewer CPU cores.
High-dimensional approximate nearest neighbor (ANN) indexes suffer from redundancy and high search overhead when serialized into generic file formats. Method: We propose eCP-FS—the first disk-based ANN indexing system that directly models the index as a human-readable, program-parsable filesystem structure, replacing conventional serialization with filesystem abstractions. Built upon the eCP indexing algorithm and standard filesystem libraries, eCP-FS enables cross-language interoperability and transparent access. Contribution/Results: Experiments show eCP-FS achieves minimal memory footprint under memory-constrained settings, with pronounced advantages in multi-index coexistence scenarios; it remains competitive even with ample memory. This work validates the feasibility and practicality of the “file-as-index” paradigm, substantially reducing debugging and maintenance complexity for ANN indexes.
This work addresses a critical limitation in existing graph-based disk indexing systems for large-scale high-dimensional vector similarity search: their performance is constrained by overlooking computational overhead, as the true bottleneck in high-dimensional settings lies in computation rather than I/O. The study is the first to reveal the intrinsic nature of this computational bottleneck and proposes a novel computation-optimized disk data layout that fully exploits modern CPU SIMD instructions. The approach integrates degree-based node caching, cluster-driven entry point selection, and an early scheduling strategy. Extensive experiments demonstrate that the proposed method significantly outperforms state-of-the-art disk-based graph index systems across multiple large-scale high-dimensional datasets, achieving performance comparable to—or even surpassing—that of in-memory indexing schemes, thereby transcending the traditional I/O-centric design paradigm.
Managing and retrieving high-dimensional vector data poses significant challenges, particularly as traditional databases fail to meet performance requirements and the need for tight integration with large language models (LLMs) intensifies. Method: This paper systematically surveys four major approximate nearest neighbor search (ANNS) paradigms—hashing, tree-based indexing, graph-based methods (e.g., HNSW), and quantization (PQ/SQ)—and integrates hybrid optimization strategies. Contribution/Results: It introduces, for the first time, a “Four-Dimensional Methodology” framework tailored for industrial deployment of vector databases, analyzing trade-offs among accuracy, latency, memory footprint, and scalability. The work constructs a structured knowledge graph covering 200+ ANNS algorithms and proposes a novel paradigm for deep synergy between vector databases and LLMs. Collectively, these contributions provide both theoretical foundations and practical guidelines for system selection, architectural design, and development of AI-native database systems.
This work addresses the inefficiency of existing vector join methods in threshold queries, which suffer from redundant index traversals and excessive distance computations. To overcome these limitations, the authors propose a unified framework that integrates three key innovations: a soft work-sharing mechanism, a merged graph index that jointly embeds query and data vectors, and an adaptive hybrid search strategy tailored for out-of-distribution queries. Evaluated on eight benchmark datasets, the proposed approach consistently outperforms state-of-the-art methods, achieving substantially lower computational overhead while maintaining high recall. The results demonstrate a superior trade-off between efficiency and accuracy for approximate threshold vector joins.
Traditional vector databases treat metadata as flat scalar attributes, which inadequately captures hierarchical directory semantics, leading to inefficient range queries, high overhead for structural updates, and challenges in maintaining consistency. This work introduces directory semantics as a first-class feature in vector databases, proposing Directory Semantic Query (DSQ) and Directory Semantic Maintenance (DSM) operations. To preserve directory topology and avoid the latency and write amplification caused by path expansion, we design TrieHI, a Trie-based hierarchical index that enables efficient recursive retrieval and low-cost structural modifications. Extensive experiments on ByteDance’s Viking engine demonstrate the superiority of our approach. We also release two large-scale datasets, WIKI-Dir and ARXIV-Dir, and have integrated TrieHI into OpenViking, an open-source context database for AI agents.
This study addresses the lack of systematic evaluation of hybrid search mechanisms that combine semantic retrieval with metadata filtering in existing vector databases. We propose a novel relevance metric, Global-Local Selectivity (GLS), construct MoReVec—the first benchmark dataset supporting filtered retrieval—and extend ANN-Benchmarks to enable unified evaluation of hybrid search performance. Through comprehensive experiments integrating diverse filtering strategies into FAISS, Milvus, and pgvector with IVFFlat and HNSW indexes, we demonstrate that engine-level algorithmic integration critically governs performance: Milvus achieves more stable recall via hybrid execution, pgvector’s optimizer often selects suboptimal query plans, and IVFFlat outperforms HNSW under low-selectivity queries. Our findings culminate in practical configuration guidelines that offer both theoretical insights and actionable recommendations for efficient hybrid search deployment.
This work addresses the challenge of balancing fine-grained semantic representation and retrieval efficiency in multi-vector retrieval, which has been hindered by the absence of efficient indexing mechanisms. To this end, we propose GEM, a native graph-based indexing framework that, for the first time, designs an index structure specifically for sets of vectors. GEM integrates set-level clustering, local proximity graph connectivity, and global navigation, while decoupling graph construction metrics from relevance scoring. It further introduces semantic shortcuts and a multi-entry beam search mechanism, enhanced with quantized distance estimation, to significantly accelerate retrieval. Experimental results demonstrate that GEM achieves up to a 16× speedup over existing methods across multiple benchmarks, while maintaining or even improving retrieval accuracy.
Existing learned indexes struggle to simultaneously achieve high concurrency, durability, and low intrusiveness under write-intensive workloads. This work proposes a hierarchical learned indexing architecture that leverages the separation between Memtables and SST files in RocksDB to enable targeted optimizations at both memory and disk layers. By reusing structural knowledge across Memtables, the approach mitigates the overhead of frequent index reconstruction, while a block-aware, read-only learned index ensures that lookups complete within a single I/O in the worst case—without requiring modifications to the storage layer or read path. Experimental results demonstrate that, across diverse large-scale workloads, the proposed method improves write throughput by up to 1.5× and read throughput by up to 2.1× compared to state-of-the-art systems.