MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the excessive memory overhead in dense retrieval caused by maintaining multiple indices for varying dimensionalities and bitrates. To overcome this limitation, we propose Matryoshka Residual Vector Quantization (MRVQ), a method that optimizes frozen embeddings through post-training processing by truncating residual stages or coordinates. This approach enables a single index to flexibly accommodate diverse dimensionality-rate combinations, achieving versatile multi-purpose indexing. Experimental results demonstrate that MRVQ reduces memory consumption by 17.8× to 22× compared to maintaining three independent indices. Although its retrieval quality is marginally lower than that of optimal single-rate configurations, it significantly outperforms conventional Product Quantization (PQ) methods. Ultimately, MRVQ facilitates efficient and elastic vector search with minimal memory footprint.
📝 Abstract
Dense-retrieval services must switch among embedding-prefix dimensions and index bit rates as latency, quality, and memory budgets change. Tuning a quantizer separately for each rate gives the best quality, but the retrieval tier then holds several code streams and quantizer states at once. We introduce Matryoshka Residual Vector Quantization (MRVQ), a post-hoc residual quantizer for frozen embeddings. Its maximum-rate code can be truncated two ways: dropping residual stages lowers the rate, and dropping embedding coordinates lowers the dimension. One resident artifact therefore serves every (dimension, rate) pair we evaluate. Across FiQA and NFCorpus, four embedding families, and {4, 8, 16}-byte codes, MRVQ is the lowest-RAM design we evaluate. It uses 17.8-22.0x less memory than three separately trained QINCo2 indices, and 1.89-2.02x less than a lean shared-model steelman. The saving is not free: per-rate QINCo2 is 0.026-0.107 nDCG@10 better on FiQA. But MRVQ beats PQ, OPQ, and AdANNS-OPQ at matched code size. We also evaluate a low-build-cost PCA-scalar design that attains quality comparable to RaBitQ and its extension while fitting 420x faster at the median. Finally, we report two negative results: QINCo2 collapses when trained at high rates, and a ranking-bound hypothesis misses its pre-specified acceptance criteria. MRVQ is therefore a low-memory operating point for elastic retrieval, not a universal quality winner.
Problem

Research questions and friction points this paper is trying to address.

vector search
residual vector quantization
elastic retrieval
memory efficiency
dense retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

Residual Vector Quantization
Elastic Vector Search
Matryoshka Representation
Memory Efficiency
Dense Retrieval
🔎 Similar Papers
2024-01-16arXiv.orgCitations: 76