Score
Design, implement, and analyze pruning rules and algorithms for inverted indexes that avoid accessing irrelevant postings, clusters, or vectors during similarity search by removing or skipping portions of inverted lists; this includes batch-filtering methods and mechanisms that achieve constant- or logarithmic-time elimination of clusters or groups. Develop and apply cosine-law-based pruning and other tight lower-bound computations to reduce unnecessary cluster and vector accesses and to maintain efficient update/query behavior for inverted-list structures.
Existing graph-based indexes (e.g., HNSW) suffer from poor connectivity, low recall, and high latency for nearest-neighbor search under complex hard filtering predicates. This paper proposes a multi-index ensemble architecture coupled with a learnable three-dimensional analytical model to enable dynamic, query-time selection of the optimal index. Our approach integrates workload-aware index packaging, parallel multi-index construction, and a lightweight HNSW variant—achieving high recall while significantly improving efficiency. Experiments show that our method accelerates queries by up to 8.06× over state-of-the-art baselines, reduces indexing time to just 1% of conventional methods, and incurs only 2.15× the memory overhead of standard HNSW. The core innovation lies in decoupling filtering logic into a collaborative multi-index mechanism and employing a learnable analytical model to drive real-time index selection—thereby overcoming the performance bottlenecks inherent to constrained graph traversal in predicate-aware ANN search.
This work addresses the high query latency in IVF-based vector retrieval caused by coarse-grained execution. It proposes CLIP, a lightweight pruning technique that, for the first time, leverages the monotonicity of the lower bound derived from the cosine law to enable O(1) inter-cluster elimination and logarithmic-time intra-cluster batch filtering, thereby supporting dual-level pruning at both cluster and vector granularities. Furthermore, integrating ideas from LSM trees, the authors design LSM-IVF to support efficient dynamic updates. Experimental results demonstrate that CLIP achieves a 78% pruning ratio and improves query efficiency by 69% in static settings, while LSM-IVF increases throughput by 141% in dynamic scenarios with manageable update overhead.
Existing vector similarity search systems face a fundamental efficiency–accuracy trade-off when supporting joint queries with attribute filtering. Method: This paper proposes a filtering-condition-centric vector indexing framework that directly embeds attribute logic into the vector space. Its core innovation is a geometric transformation ψ(v, f, α), provably guaranteeing retrieval accuracy and enabling plug-and-play filtering enhancement for mainstream ANN indexes—including HNSW, FAISS, and ANNOY—without modifying their internal structures. The transformation exhibits theoretical robustness against data distribution shifts. A meta-index architecture and theory-driven filtering encoding further ensure scalability and stability. Results: The method matches state-of-the-art recall while achieving 2.6–3.0× higher throughput; it maintains consistent performance across multi-filter scenarios and distributional variations, significantly outperforming both decoupled and jointly optimized alternatives.
To address the low online top-k document retrieval efficiency in learned sparse retrieval, this paper proposes Dynamic Superblock Pruning (DSP). DSP introduces a hierarchical superblock structure—departing from conventional flat blocks or clusters—and models the hierarchical relationships among document blocks to enable early, group-level pruning at the superblock granularity. It further integrates a dynamic threshold-driven selection and pruning strategy, ensuring rank-safety or controllable approximation accuracy under high-relevance competition constraints. The method is fully optimized for single-threaded CPU execution, requiring no specialized hardware. Evaluated on the MS MARCO passage collection, DSP significantly outperforms state-of-the-art sparse retrieval baselines: it achieves substantial speedup in online retrieval while maintaining high recall.
To address the longstanding trade-off between high compression ratio and fast decoding in inverted index compression, this paper proposes a novel integer list compression method based on ternary (trit) encoding and context modeling. The core innovation lies in the first application of context-adaptive arithmetic coding to model ternary delta sequences, coupled with a lightweight, inverted-index-specific context modeling mechanism designed to exploit the characteristic skewed distribution of posting list gaps. Experimental evaluation across multiple standard benchmarks demonstrates that the proposed method consistently achieves superior compression ratios compared to Binary Interpolative Coding, while simultaneously delivering significant decoding speedups. All source code and experimental data are publicly released, establishing a new paradigm for efficient inverted index compression.
This work addresses the performance bottleneck in lakehouse architectures caused by the separation of structured filtering and approximate nearest neighbor (ANN) search. It proposes an efficient co-design that embeds IVF vector indexes into Parquet file footers, enabling file-level ANN search on Apache Iceberg tables after leveraging existing file-pruning mechanisms—such as partition pruning and zone maps—to first eliminate irrelevant data files. This approach achieves, for the first time in open lakehouse formats, joint filtering and vector retrieval without modifying the query engine, supports distributed and non-intrusive index construction, and incorporates consistent hashing-based caching to mitigate object storage latency. Experiments show that on an 11.5M × 768 dataset, ANN search with selective filtering is 32× faster than brute-force search while maintaining recall@10 ≥ 0.90; on 5M IBM Granite embeddings, post-join filtered queries accelerate by 94×, reducing latency from 14.7 seconds to 157 milliseconds.
This study addresses the inefficiency of traditional data skipping techniques when applied to database filters based on black-box machine learning models. To bridge this gap, the work introduces the first extension of data skipping to ML-based filtering, proposing a lightweight augmented metadata structure grounded in bounded two-dimensional convex hulls. This structure synergistically integrates Parquet’s native min-max statistics, ML query semantics, and neural network verification techniques—including ReLU architecture analysis—to achieve substantially enhanced pruning efficacy with minimal storage overhead. Experimental evaluation on TPC-H and TPC-DS benchmarks under low-selectivity (<0.1%) query workloads demonstrates that the approach improves average pruning rates from 27.4% to 38.31% and accelerates end-to-end query execution by 1.07× compared to a PyTorch-integrated DuckDB baseline.
This work addresses the limitation of existing vector retrieval systems, which support only simple numerical or categorical constraints and struggle with complex graph-structured filtering requirements. To bridge this gap, the authors propose DLH, the first method to integrate graph-range filtering into approximate nearest neighbor (ANN) search. DLH constructs distance-aware label sets and compresses them into Bloom filters to enable efficient graph-range filtering. Furthermore, it introduces a query-node hash index caching mechanism (DLH-M) that exploits query locality to accelerate retrieval. Experimental results demonstrate that DLH achieves up to a 70.3% increase in throughput across multiple datasets while maintaining recall above 98.5%, with only modest additional storage overhead.