🤖 AI Summary
This paper addresses lossless compression of numerical sequences, proposing the Pcodec format and algorithm. Methodologically: (1) it introduces an entropy-estimation-driven adaptive binning algorithm that achieves rapid convergence to the true entropy—proven theoretically to incur only $O(1/k)$ bits of redundancy; and (2) it establishes a two-stage preprocessing framework combining latent-variable pattern decomposition with differential encoding, jointly leveraging SIID (Stationary, Independent, Identically Distributed) modeling to enhance accuracy. The key contribution is a unified optimization of compression ratio and speed: on six real-world datasets, Pcodec improves compression ratios by 29%–94% while simultaneously reducing compression time. To our knowledge, this work presents the first end-to-end solution for numerical sequence compression that provides both rigorous theoretical guarantees and competitive practical performance.
📝 Abstract
We present Pcodec (Pco), a format and algorithm for losslessly compressing numerical sequences. Pco's core and most novel component is a binning algorithm that quickly converges to the true entropy of smoothly, independently, and identically distributed (SIID) data. To automatically handle more general data, Pco has two opinionated preprocessing steps. The first step, Pco's mode, decomposes the data into more smoothly distributed latent variables. The second step, delta encoding, makes the latents more independently and identically distributed. We prove that, given $k$ bins, binning uses only $mathcal{O}(1/k)$ bits more than the SIID data's entropy. Additionally, we demonstrate that Pco achieves 29-94% higher compression ratio than other approaches on six real-world datasets while using less compression time.