🤖 AI Summary
To address the longstanding trade-off between high compression ratio and fast decoding in inverted index compression, this paper proposes a novel integer list compression method based on ternary (trit) encoding and context modeling. The core innovation lies in the first application of context-adaptive arithmetic coding to model ternary delta sequences, coupled with a lightweight, inverted-index-specific context modeling mechanism designed to exploit the characteristic skewed distribution of posting list gaps. Experimental evaluation across multiple standard benchmarks demonstrates that the proposed method consistently achieves superior compression ratios compared to Binary Interpolative Coding, while simultaneously delivering significant decoding speedups. All source code and experimental data are publicly released, establishing a new paradigm for efficient inverted index compression.
📝 Abstract
Inverted indexes allow to query large databases without needing to search in the database at each query. An important line of research is to construct the most efficient inverted indexes, both in terms of compression ratio and time efficiency. In this article, we show how to use trit encoding, combined with contextual methods for computing inverted indexes. We perform an extensive study of different variants of these methods and show that our method consistently outperforms the Binary Interpolative Method -- which is one of the golden standards in this topic -- with respect to compression size. We apply our methods to a variety of datasets and make available the source code that produced the results, together with all our datasets.