Compact and Efficient Indexes for Learned Sparse Retrieval

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high memory footprint of learned sparse retrieval indices by proposing SEISMIC, a pluggable architectural optimization framework compatible with systems such as KANNOLO. Methodologically, SEISMIC employs medoid-based metadata reduction for inverted indices and achieves forward index compression through lexical reordering, DOTPACKING 8-bit packing, and 4-bit codebook quantization. Furthermore, it introduces JUMPDOT, a block dot-product kernel that accelerates sparse computations. Evaluated on the MS MARCO benchmark, SEISMIC delivers a 5.3× speedup in retrieval latency and a 3× reduction in memory consumption while maintaining equivalent accuracy. Even under severely resource-constrained settings, it sustains a 1.9× speedup alongside a 3.9× memory reduction, substantially improving the speed–space trade-off in learned sparse retrieval.
📝 Abstract
This paper investigates how to substantially reduce the memory footprint of learned sparse retrieval indexes without sacrificing the efficiency of state-of-the-art retrieval data structures. Building on SEISMIC, we revisit both levels of its design: the inverted index used to select candidates and the forward index used to score them. For the inverted index, we replace costly per-block summaries with medoids, namely existing documents elected as block representatives, collapsing the per-block metadata from a sparse vector to a single document identifier. For the forward index, we compress both components and values. We reorder the vocabulary to place co-occurring components closer together and encode the resulting $Δ$-gaps with DOTPACKING8, a SIMD-friendly bit-packing scheme that fuses decompression with dot-product evaluation; values are quantized with compact per-component 4-bit codebooks fitted to each component's distribution. We further introduce JUMPDOT, a blocked dot-product kernel tailored for queries that contain only a few non-zero entries. Our forward-index compression is independent of SEISMIC and can be plugged into any system relying on forward-index-based scoring, as we demonstrate by integrating it into KANNOLO. A comprehensive evaluation on MS MARCO with three state-of-the-art learned sparse encoders shows that our solutions markedly improve the speed-space trade-off of learned sparse retrieval: at equal accuracy, our indexes answer queries up to 5.3x faster than the best competitor while using about 3x less memory, and in the most memory-constrained regime, they remain up to 1.9x faster while using up to 3.9x less memory.
Problem

Research questions and friction points this paper is trying to address.

Learned Sparse Retrieval
Index Compression
Memory Footprint
Inverted Index
Forward Index
Innovation

Methods, ideas, or system contributions that make the work stand out.

Learned Sparse Retrieval
Index Compression
SIMD Bit-packing
Quantization
Dot-product Kernel
🔎 Similar Papers
No similar papers found.