indexing strategies

Designs and implements index structures and complete indexing systems for efficient storage, lookup, and retrieval across modalities (e.g., full-text, hash-based, spatial, graph, and vector/similarity search), producing compact searchable representations and scalable I/O. Builds and analyzes index construction and update pipelines—including incremental updates, ANN/vector index construction, semantic and similarity indexing—and optimizes indexing algorithms and structures for lookup, retrieval latency, throughput, and resource trade-offs.

indexingstrategies

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
2.7
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Accelerating Graph Indexing for ANNS on Modern CPUs

Feb 25, 2025
MW
Mengzhao Wang
🏛️ Zhejiang University | Alibaba Group

To address the prohibitively long indexing time and poor CPU-architecture compatibility of graph-based indexes in high-dimensional approximate nearest neighbor search (ANNS), this paper proposes Flash—a hardware-aware compact encoding strategy. Flash innovatively integrates vector quantization with key CPU architectural features, including SIMD parallelism and cache locality, enabling efficient distance computation while maintaining bounded quantization error. By jointly optimizing compact encoding, memory access patterns, and cache-friendly graph construction, Flash achieves 10.4×–22.9× speedup in index construction across eight real-world datasets ranging from 10 million to one billion vectors. Crucially, this acceleration comes without sacrificing retrieval accuracy or query latency—indeed, both are preserved or improved. Flash thus bridges the gap between algorithmic efficiency and modern hardware utilization in large-scale ANNS.

Distance computation dominates indexing timeHigh indexing time in graph-based ANNS methodsOptimizing graph indexing for modern CPU architectures

This work addresses the challenge of balancing recall accuracy and throughput efficiency under frequent vector updates in the era of large language models. The authors propose Yi, a novel graph-based vector indexing system that pioneers support for in-place updates. Guided by a “decompose-to-integrate” design philosophy, Yi introduces a vector-level update mechanism that overcomes the bottlenecks of traditional batch merging. It integrates three core components—a tasklet execution engine, an asynchronous buffer manager, and a vector file system—to enable efficient online updates and high-quality retrieval. Experiments on a dataset of 800 million vectors demonstrate that Yi achieves 1.75× higher update throughput, 1.8× greater concurrent search throughput, and reduces peak memory usage to 73% of baseline systems, delivering superior performance even with fewer CPU cores.

graph-based indexin-place updatessearch recall

High-dimensional approximate nearest neighbor (ANN) indexes suffer from redundancy and high search overhead when serialized into generic file formats. Method: We propose eCP-FS—the first disk-based ANN indexing system that directly models the index as a human-readable, program-parsable filesystem structure, replacing conventional serialization with filesystem abstractions. Built upon the eCP indexing algorithm and standard filesystem libraries, eCP-FS enables cross-language interoperability and transparent access. Contribution/Results: Experiments show eCP-FS achieves minimal memory footprint under memory-constrained settings, with pronounced advantages in multi-index coexistence scenarios; it remains competitive even with ample memory. This work validates the feasibility and practicality of the “file-as-index” paradigm, substantially reducing debugging and maintenance complexity for ANN indexes.

Assesses suitability of eCP-FS for memory-constrained environmentsCompares eCP-FS with memory-based and disk-based ANN indexesEvaluates performance penalty of file-based ANN indexes

This work addresses a critical limitation in existing graph-based disk indexing systems for large-scale high-dimensional vector similarity search: their performance is constrained by overlooking computational overhead, as the true bottleneck in high-dimensional settings lies in computation rather than I/O. The study is the first to reveal the intrinsic nature of this computational bottleneck and proposes a novel computation-optimized disk data layout that fully exploits modern CPU SIMD instructions. The approach integrates degree-based node caching, cluster-driven entry point selection, and an early scheduling strategy. Extensive experiments demonstrate that the proposed method significantly outperforms state-of-the-art disk-based graph index systems across multiple large-scale high-dimensional datasets, achieving performance comparable to—or even surpassing—that of in-memory indexing schemes, thereby transcending the traditional I/O-centric design paradigm.

compute-boundgraph-based indexhigh-dimensional vectors

A Comprehensive Survey on Vector Database: Storage and Retrieval Technique, Challenge

Oct 18, 2023
YH
Yikun Han
🏛️ University of Michigan | Chinese Academy of Sciences

Managing and retrieving high-dimensional vector data poses significant challenges, particularly as traditional databases fail to meet performance requirements and the need for tight integration with large language models (LLMs) intensifies. Method: This paper systematically surveys four major approximate nearest neighbor search (ANNS) paradigms—hashing, tree-based indexing, graph-based methods (e.g., HNSW), and quantization (PQ/SQ)—and integrates hybrid optimization strategies. Contribution/Results: It introduces, for the first time, a “Four-Dimensional Methodology” framework tailored for industrial deployment of vector databases, analyzing trade-offs among accuracy, latency, memory footprint, and scalability. The work constructs a structured knowledge graph covering 200+ ANNS algorithms and proposes a novel paradigm for deep synergy between vector databases and LLMs. Collectively, these contributions provide both theoretical foundations and practical guidelines for system selection, architectural design, and development of AI-native database systems.

Compare advanced VDB solutions with strengths and limitationsExplore coupling VDBs with large language modelsReview storage and retrieval techniques in vector databases

Latest Papers

What's happening recently
View more

This work addresses the inefficiency of existing vector join methods in threshold queries, which suffer from redundant index traversals and excessive distance computations. To overcome these limitations, the authors propose a unified framework that integrates three key innovations: a soft work-sharing mechanism, a merged graph index that jointly embeds query and data vectors, and an adaptive hybrid search strategy tailored for out-of-distribution queries. Evaluated on eight benchmark datasets, the proposed approach consistently outperforms state-of-the-art methods, achieving substantially lower computational overhead while maintaining high recall. The results demonstrate a superior trade-off between efficiency and accuracy for approximate threshold vector joins.

distance computationindex traversalthreshold-based

Traditional vector databases treat metadata as flat scalar attributes, which inadequately captures hierarchical directory semantics, leading to inefficient range queries, high overhead for structural updates, and challenges in maintaining consistency. This work introduces directory semantics as a first-class feature in vector databases, proposing Directory Semantic Query (DSQ) and Directory Semantic Maintenance (DSM) operations. To preserve directory topology and avoid the latency and write amplification caused by path expansion, we design TrieHI, a Trie-based hierarchical index that enables efficient recursive retrieval and low-cost structural modifications. Extensive experiments on ByteDance’s Viking engine demonstrate the superiority of our approach. We also release two large-scale datasets, WIKI-Dir and ARXIV-Dir, and have integrated TrieHI into OpenViking, an open-source context database for AI agents.

directory semanticshierarchical metadatarecursive retrieval

This study addresses the lack of systematic evaluation of hybrid search mechanisms that combine semantic retrieval with metadata filtering in existing vector databases. We propose a novel relevance metric, Global-Local Selectivity (GLS), construct MoReVec—the first benchmark dataset supporting filtered retrieval—and extend ANN-Benchmarks to enable unified evaluation of hybrid search performance. Through comprehensive experiments integrating diverse filtering strategies into FAISS, Milvus, and pgvector with IVFFlat and HNSW indexes, we demonstrate that engine-level algorithmic integration critically governs performance: Milvus achieves more stable recall via hybrid execution, pgvector’s optimizer often selects suboptimal query plans, and IVFFlat outperforms HNSW under low-selectivity queries. Our findings culminate in practical configuration guidelines that offer both theoretical insights and actionable recommendations for efficient hybrid search deployment.

Filtered Approximate Nearest Neighbor SearchFiltering StrategiesHybrid Search

This work addresses the challenge of balancing fine-grained semantic representation and retrieval efficiency in multi-vector retrieval, which has been hindered by the absence of efficient indexing mechanisms. To this end, we propose GEM, a native graph-based indexing framework that, for the first time, designs an index structure specifically for sets of vectors. GEM integrates set-level clustering, local proximity graph connectivity, and global navigation, while decoupling graph construction metrics from relevance scoring. It further introduces semantic shortcuts and a multi-entry beam search mechanism, enhanced with quantized distance estimation, to significantly accelerate retrieval. Experimental results demonstrate that GEM achieves up to a 16× speedup over existing methods across multiple benchmarks, while maintaining or even improving retrieval accuracy.

high-dimensional vectorsindexing algorithmsmulti-vector retrieval

Existing learned indexes struggle to simultaneously achieve high concurrency, durability, and low intrusiveness under write-intensive workloads. This work proposes a hierarchical learned indexing architecture that leverages the separation between Memtables and SST files in RocksDB to enable targeted optimizations at both memory and disk layers. By reusing structural knowledge across Memtables, the approach mitigates the overhead of frequent index reconstruction, while a block-aware, read-only learned index ensures that lookups complete within a single I/O in the worst case—without requiring modifications to the storage layer or read path. Experimental results demonstrate that, across diverse large-scale workloads, the proposed method improves write throughput by up to 1.5× and read throughput by up to 2.1× compared to state-of-the-art systems.

learned indexingminimal system modificationproduction database integration

Hot Scholars

TG

Travis Gagie

Associate Professor at Dalhousie University
data structuresdata compression
TP

Themis Palpanas

Distinguished Professor, University Paris Cite, French University Institute (IUF)
data managementdata sciencedata/time seriesanomaly detection
DK

Dominik Köppl

Faculty for Engineering, University of Yamanashi
stringologyalgorithms and data structurescombinatorics on words
YF

Yixiang Fang

Associate Professor, The Chinese University of Hong Kong, Shenzhen
Data managementdata miningand artificial intelligence
SP

Solon P. Pissis

Senior Researcher, CWI
AlgorithmsData StructuresBig DataData Mining