Pcodec: Better Compression for Numerical Sequences

📅 2025-02-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses lossless compression of numerical sequences, proposing the Pcodec format and algorithm. Methodologically: (1) it introduces an entropy-estimation-driven adaptive binning algorithm that achieves rapid convergence to the true entropy—proven theoretically to incur only $O(1/k)$ bits of redundancy; and (2) it establishes a two-stage preprocessing framework combining latent-variable pattern decomposition with differential encoding, jointly leveraging SIID (Stationary, Independent, Identically Distributed) modeling to enhance accuracy. The key contribution is a unified optimization of compression ratio and speed: on six real-world datasets, Pcodec improves compression ratios by 29%–94% while simultaneously reducing compression time. To our knowledge, this work presents the first end-to-end solution for numerical sequence compression that provides both rigorous theoretical guarantees and competitive practical performance.

Technology Category

Data Mining & Knowledge Management: Data CompressionCognitive Modeling & Cognitive Systems: Neural Spike CodingMachine Learning: Learning on the Edge & Model Compression

Application Category

Graph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
We present Pcodec (Pco), a format and algorithm for losslessly compressing numerical sequences. Pco's core and most novel component is a binning algorithm that quickly converges to the true entropy of smoothly, independently, and identically distributed (SIID) data. To automatically handle more general data, Pco has two opinionated preprocessing steps. The first step, Pco's mode, decomposes the data into more smoothly distributed latent variables. The second step, delta encoding, makes the latents more independently and identically distributed. We prove that, given $k$ bins, binning uses only $mathcal{O}(1/k)$ bits more than the SIID data's entropy. Additionally, we demonstrate that Pco achieves 29-94% higher compression ratio than other approaches on six real-world datasets while using less compression time.
Problem

Research questions and friction points this paper is trying to address.

Losslessly compress numerical sequences efficiently
Handle general data with preprocessing steps
Achieve higher compression ratios than existing methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

Binning algorithm for entropy convergence
Mode decomposition for smooth distribution
Delta encoding for SIID transformation