Score
Designs and evaluates algorithms, codecs, and tools that reduce storage or transmission cost of discrete sequences and state-space representations, including methods for sequence compression, state-space compression, and streaming decompression. Work includes developing and tuning compression parameters, implementing compression/decompression pipelines, and analyzing compression ratio, latency, and fidelity across data-compression techniques.
To address the need for efficient lossless compression of time-series data, this paper proposes a two-stage unified framework comprising data transformation (e.g., differencing, predictive coding) followed by entropy coding (e.g., Huffman, arithmetic coding). We conduct the largest end-to-end lossless compression benchmark to date, evaluating combinations of transformation and coding methods across synthetic and diverse real-world time-series datasets. A standardized evaluation protocol and systematic ablation analysis are introduced to quantify individual component contributions and reveal the sensitivity of algorithm performance to intrinsic data characteristics. Key contributions include: (i) empirical validation that holistic pipeline design—not isolated components—dominates overall compression efficacy; (ii) identification of critical synergies among transformation and coding modules; and (iii) establishment of principled guidelines for scenario-aware algorithm selection and customized pipeline construction, bridging theoretical insight with practical deployment.
This work addresses the high latency introduced by traditional decompression in scientific data analysis, which undermines the storage and transmission benefits of compression. To overcome this limitation, the authors propose a multi-stage, error-bounded decompression and homomorphic analysis framework. By abstracting a generic compression pipeline, the framework enables hierarchical partial decompression and introduces homomorphic operation algorithms tailored to three representative scientific analysis tasks, allowing computations to be performed directly on intermediate compressed representations without full decompression. Implemented atop four mainstream compressors and evaluated across five real-world datasets, the approach consistently reduces data access latency and significantly improves analytical efficiency across diverse workloads.
Lossless compression algorithms face inherent trade-offs among multidimensional performance metrics—particularly compression ratio and encoding/decoding speed—posing challenges in latency-sensitive, high-fidelity applications such as medical imaging. Method: This paper introduces the first unified, quantifiable multi-objective evaluation framework for lossless compression, employing normalized weighted modeling to dynamically balance compression ratio and speed across diverse data modalities (e.g., images, text). Contribution/Results: Its key innovation lies in a standardized multi-objective scoring model that bridges the gap between theoretical metrics and practical deployment requirements. Extensive experiments demonstrate the framework’s robustness in identifying scenario-optimal compressors: learned codecs achieve superior compression ratios, while traditional algorithms retain advantages in speed-critical tasks. The framework enables principled, application-aware algorithm selection without requiring domain-specific re-engineering.
Addressing the challenge of simultaneously achieving progressive decompression, random access, and high compression quality/speed in lossy scientific data compression, this paper proposes the first unified framework supporting both decompression modes. Our method employs hierarchical data partitioning and hierarchical prediction, integrated with error-bounded lossy compression and streaming-based codec design. It maintains reconstruction accuracy comparable to state-of-the-art non-streaming compressors (e.g., SZ3), while significantly improving decompression throughput—up to 6.7× faster than SZ3. To our knowledge, this is the first approach to jointly optimize high fidelity, high throughput, and flexible data access. The framework enables practical online analysis and on-demand data retrieval, bridging a critical gap between theoretical compression efficiency and real-world scientific computing workflows.
The field of bounded-lossy compression for scientific data lacks a systematic survey and unified classification framework. Method: This work proposes the first six-dimensional taxonomy, categorizing existing approaches into six model classes; systematically analyzes 46 state-of-the-art compressors—elucidating their design principles, error-control mechanisms, and application domains—and distills five core technical components: predictive coding, transform coding, quantization, entropy coding, and parallel/distributed architectures. Furthermore, it establishes domain-specific compression design methodologies and selection guidelines tailored to high-performance computing (HPC), climate modeling, and particle physics. Contribution/Results: The framework enables high-fidelity, high-ratio scientific data management and has become a benchmark reference across multiple disciplines.
This study addresses the optimization of general-purpose lossless compression under realistic resource constraints—specifically, memory usage capped at 8 GB and decompressor size limited to 1 MB—by organizing an international challenge based on a public training set and a hidden test set comprising 16 heterogeneous files. Performance is evaluated multidimensionally using compression ratio, compression/decompression time, Weissman score, and Pareto front analysis. The generalization capability of submissions is further assessed on external large-scale datasets, while Normalized Compression Distance (NCD) is employed to analyze inter-solution relationships. Among 117 valid submissions, several outperformed mainstream tools on external data, underscoring the critical role of advanced probabilistic modeling and effective multi-objective trade-offs in enhancing compression performance.
General-purpose compression algorithms often struggle to simultaneously achieve high compression ratios and high throughput with low overhead, whereas specialized compressors, while offering superior performance, incur high development and maintenance costs and suffer from limited applicability. This work proposes a novel “graph-based” compression framework that, for the first time, models the compression process as a modular composition of encoders and decoders represented by a directed acyclic graph. By integrating a self-describing format with a universal decoder (OpenZL), the framework unifies the generality of generic methods with the performance of specialized ones. The approach substantially reduces the cost of developing and deploying domain-specific compressors, outperforming mainstream general-purpose compressors in both compression ratio and speed across multiple real-world datasets. It remains competitive with deep learning–based methods while operating orders of magnitude faster, and internal adoption at Meta has reduced development cycles from months to days.
This study addresses the trade-off between compression efficiency and transmission delay when using variable-length coding for real-time text streaming over fixed-rate channels, where queueing delays arise from fluctuating codeword lengths. For the first time, it systematically evaluates predictive coding driven by large language models in this context, employing GPT-2 (124M) and Llama 3.2 (3B) as causal predictors combined with entropy coding schemes including Shannon, Huffman, arithmetic coding, rANS, and gzip. Experimental results demonstrate that increasing model size by a factor of 25 reduces bits per character by approximately 38%. Among entropy coders, Huffman coding proves more practical in over-provisioned channels due to its zero algorithmic delay, whereas arithmetic coding approaches the theoretical compression limit at the cost of introducing decoding latency.
Traditional three-stage Huffman coding incurs significant computational and latency overhead in multi-accelerator communication due to the need for real-time frequency analysis, codebook generation, and transmission, making it ill-suited for low-latency requirements. This work proposes a single-stage Huffman encoder that constructs a fixed codebook based on the average probability distribution derived from historical data batches, eliminating the need for real-time codebook generation and transmission. For the first time, this approach is applied to lossless compression of cross-layer and sharded tensors. Evaluated on the Gemma 2B model, the method achieves a compression ratio only 0.5% lower than per-shard Huffman coding and lies within 1% of the Shannon limit, substantially improving communication efficiency while closely approaching the theoretical compression bound.
In industrial and IoT applications, high-ratio compression of real-time/historical process data reduces storage costs and improves efficiency but risks degrading the accuracy of statistical analysis, anomaly detection, and machine learning models. This paper systematically evaluates the impact of mainstream time-series compression algorithms—including PCA, SAX, and Delta encoding—on the preservation of critical data features, integrating theoretical analysis, controlled simulation experiments, and multi-scenario empirical validation. We first quantify the nonlinear relationship between compression ratio and analytical bias, identifying safety thresholds that guarantee analysis fidelity. Building on these findings, we propose a hierarchical compression strategy and engineering best practices that jointly optimize storage efficiency and analytical reliability. Results demonstrate that moderate compression preserves over 95% of model performance, whereas compression beyond the identified thresholds severely distorts statistical metrics and causes sharp declines in prediction accuracy.