🤖 AI Summary
This work proposes Split-RaBitQ, a framework addressing the substantial SSD read overhead and I/O latency-hiding challenges in large-scale vector retrieval. The method introduces sub-bit memory representations that retain partial binary codes in DRAM to enable efficient candidate pruning. To preserve accuracy, it designs an unbiased distance estimator with probabilistic error bounds. Furthermore, an asynchronous search pipeline built upon an IVF-RaBitQ extension orchestrates computation and I/O efficiently while providing fine-grained SSD access control. Experimental results demonstrate that Split-RaBitQ improves system throughput by 1.74× and reduces SSD read volume by 3.8×, supporting high-concurrency queries at billion-scale on a single machine.
📝 Abstract
SSD-resident approximate nearest-neighbor search is essential when vector collections exceed DRAM capacity. The challenge is to reduce SSD reads and hide I/O latency through concurrent reads and overlap with computation. However, for graph-based search, progressive candidate discovery limits advance I/O planning while IVF search reads inverted lists in full, incurring unnecessary SSD reads. In this paper, we present RaBitQ-SSD, an extension of IVF-RaBitQ for SSD-resident vector search. Specifically, we propose Split-RaBitQ, which retains a configurable prefix of each binary code in DRAM, allowing in-memory code storage below one bit per dimension. Using this partial representation, it provides an unbiased distance estimator with a probabilistic error bound. Once the in-memory coarse quantizer identifies candidate lists, this bound supports pruning individual candidates within them, enabling finer-grained SSD access. We also design an asynchronous search pipeline that coordinates candidate pruning with SSD read scheduling to reduce unnecessary reads while overlapping I/O with computation. On datasets ranging from $5$ million to $1$ billion vectors, RaBitQ-SSD delivers up to $1.74\times$ the throughput of graph-based baselines at $90\%$ recall, while reducing SSD page reads by up to $3.8\times$. Its on-SSD indexes are up to $7.0\times$ smaller and $10.0\times$ faster to build than DiskANN. We also build and search an index of $10$ billion vectors on a single machine with one SSD, achieving over $3{,}000$ queries per second at $90\%$ recall.