Score
Designs, implements, and evaluates algorithms, encoders/decoders, and file or stream formats that reduce the storage or transmission size of digital data using lossless or lossy techniques. This work includes source modeling, transforms and quantization, entropy coding, rate–distortion tradeoffs, and practical considerations such as computational cost, memory use, and robustness of compression and decompression.
Lossless compression algorithms face inherent trade-offs among multidimensional performance metrics—particularly compression ratio and encoding/decoding speed—posing challenges in latency-sensitive, high-fidelity applications such as medical imaging. Method: This paper introduces the first unified, quantifiable multi-objective evaluation framework for lossless compression, employing normalized weighted modeling to dynamically balance compression ratio and speed across diverse data modalities (e.g., images, text). Contribution/Results: Its key innovation lies in a standardized multi-objective scoring model that bridges the gap between theoretical metrics and practical deployment requirements. Extensive experiments demonstrate the framework’s robustness in identifying scenario-optimal compressors: learned codecs achieve superior compression ratios, while traditional algorithms retain advantages in speed-critical tasks. The framework enables principled, application-aware algorithm selection without requiring domain-specific re-engineering.
Exponential growth in scientific data has outpaced network bandwidth, storage capacity, and analytical capabilities, necessitating lossy compression. However, existing research lacks application-specific fidelity requirements aligned with scientific discovery goals, leading to a disconnect between algorithm design and real-world needs. Method: We conduct a systematic survey of nine representative scientific domains—including climate modeling, combustion, and cosmology—to identify cross-cutting quality constraints (e.g., error bounds on key physical quantities), compression ratios, and throughput requirements. We then develop a “discovery-oriented” compression evaluation framework and analyze error-control mechanisms and applicability boundaries of mainstream tools (SZ, ZFP, MGARD). Contribution/Results: This work delivers the first multi-disciplinary white paper on lossy compression requirements for scientific data, enabling verifiable benchmarking and co-evolution of high-fidelity, high-performance compression toolchains.
To address 6G’s demand for ultra-high compression ratios, conventional lossless compression methods have nearly reached their theoretical limits. This paper introduces LMCompress, the first framework to directly leverage large language models (LLMs) for universal lossless compression. It formulates compression as optimal induction over data distributions—approximating the incomputable Solomonoff induction. Methodologically, LMCompress integrates LLM-driven sequence modeling and probabilistic prediction, context-aware entropy coding, and model distillation with quantization for acceleration. Experiments demonstrate that LMCompress consistently outperforms state-of-the-art (SOTA) methods across diverse modalities: achieving ~2× higher compression ratios on JPEG-XL (images) and FLAC (audio), and ~4× improvement on bz2 (text), while also surpassing H.264 (video). These gains break classical information-theoretic bottlenecks and establish a novel lossless compression paradigm grounded in universal data understanding.
This study addresses the optimization of general-purpose lossless compression under realistic resource constraints—specifically, memory usage capped at 8 GB and decompressor size limited to 1 MB—by organizing an international challenge based on a public training set and a hidden test set comprising 16 heterogeneous files. Performance is evaluated multidimensionally using compression ratio, compression/decompression time, Weissman score, and Pareto front analysis. The generalization capability of submissions is further assessed on external large-scale datasets, while Normalized Compression Distance (NCD) is employed to analyze inter-solution relationships. Among 117 valid submissions, several outperformed mainstream tools on external data, underscoring the critical role of advanced probabilistic modeling and effective multi-objective trade-offs in enhancing compression performance.
This work addresses the lack of theoretical formalization and systematic implementation for “lossless” model compression. We propose LLC, the first general theoretical framework for provably lossless compression. Methodologically: (i) leveraging total differentials, we rigorously bound compression error and formally define the *lossless compression neighborhood* and higher-order analytical error bounds; (ii) we formulate quantization as a grouped knapsack problem, jointly optimizing layer-wise low-rank structures and quantization bit-widths to automatically determine the optimal lossless compression configuration. Rigorous evaluation across diverse architectures (ViT, ResNet) and datasets (ImageNet, CIFAR) confirms zero accuracy degradation post-compression, 1.8–3.2× inference speedup, and 40–65% memory reduction—without fine-tuning or heuristic design. Our core contribution is a theoretically grounded, computationally tractable, and broadly generalizable framework for lossless model compression.
General-purpose compression algorithms often struggle to simultaneously achieve high compression ratios and high throughput with low overhead, whereas specialized compressors, while offering superior performance, incur high development and maintenance costs and suffer from limited applicability. This work proposes a novel “graph-based” compression framework that, for the first time, models the compression process as a modular composition of encoders and decoders represented by a directed acyclic graph. By integrating a self-describing format with a universal decoder (OpenZL), the framework unifies the generality of generic methods with the performance of specialized ones. The approach substantially reduces the cost of developing and deploying domain-specific compressors, outperforming mainstream general-purpose compressors in both compression ratio and speed across multiple real-world datasets. It remains competitive with deep learning–based methods while operating orders of magnitude faster, and internal adoption at Meta has reduced development cycles from months to days.
This work addresses the high latency introduced by traditional decompression in scientific data analysis, which undermines the storage and transmission benefits of compression. To overcome this limitation, the authors propose a multi-stage, error-bounded decompression and homomorphic analysis framework. By abstracting a generic compression pipeline, the framework enables hierarchical partial decompression and introduces homomorphic operation algorithms tailored to three representative scientific analysis tasks, allowing computations to be performed directly on intermediate compressed representations without full decompression. Implemented atop four mainstream compressors and evaluated across five real-world datasets, the approach consistently reduces data access latency and significantly improves analytical efficiency across diverse workloads.
This work addresses the gap between Shannon’s rate-distortion theory and practical performance at finite blocklengths, focusing on the Bernoulli source under Hamming distortion. Starting from first principles, it rigorously derives the rate-distortion function \( R(D) = H(p) - H(D) \). By integrating the Blahut–Arimoto algorithm with finite-blocklength asymptotic analysis, the paper constructs a tutorial-style theoretical framework that explicitly introduces the rate-distortion dispersion \( V(D) \) to characterize the \( O(1/\sqrt{n}) \) convergence rate to the asymptotic limit. Accompanying reproducible Python simulations validate the theoretical predictions, offering both precise quantification of finite-length compression limits and a practical toolkit for applied analysis.
This work proposes DeXOR, a novel compression framework addressing the inefficiency of existing methods in handling high-precision or non-smooth streaming floating-point data. DeXOR introduces XOR-based encoding directly in decimal space for the first time, leveraging longest common prefix and suffix extraction to eliminate redundancy. The framework integrates several optimization strategies—including error-tolerant rounding, scaled truncation, exponent outlier handling, and bit-level management—to enhance compression while preserving decompression accuracy. Evaluated on 22 real-world datasets, DeXOR achieves an average 15% improvement in compression ratio and a 20% increase in decompression speed, demonstrating robustness and efficiency even under extreme conditions.