Score
Designs, builds, and evaluates algorithms and models that produce or exploit compact, compression-aware representations and compressed models while preserving information needed for downstream tasks. This includes designing lossy and lossless compression methods for embeddings and signals, compression-aware pretraining and contrastive objectives using compressed echoes, compact autoencoder architectures, hardware-aware model compression and optimization, sensitivity-profile guided compression, and metrics/procedures for compression evaluation.
Lossless compression algorithms face inherent trade-offs among multidimensional performance metrics—particularly compression ratio and encoding/decoding speed—posing challenges in latency-sensitive, high-fidelity applications such as medical imaging. Method: This paper introduces the first unified, quantifiable multi-objective evaluation framework for lossless compression, employing normalized weighted modeling to dynamically balance compression ratio and speed across diverse data modalities (e.g., images, text). Contribution/Results: Its key innovation lies in a standardized multi-objective scoring model that bridges the gap between theoretical metrics and practical deployment requirements. Extensive experiments demonstrate the framework’s robustness in identifying scenario-optimal compressors: learned codecs achieve superior compression ratios, while traditional algorithms retain advantages in speed-critical tasks. The framework enables principled, application-aware algorithm selection without requiring domain-specific re-engineering.
Existing lossless matrix compression methods fail to effectively capture structural redundancies introduced during data cleaning, augmentation, and feature engineering, leading to suboptimal efficiency in data-centric ML pipelines. This paper introduces BWARE, the first framework to deeply embed lossless compression within the outer loop of data engineering—enabling end-to-end co-design of compression and feature transformation. Its key contributions are: (1) a workload-aware compression mechanism supporting lightweight, on-the-fly morphing without decompression; (2) a column correlation- and sparsity-aware morphing mapping; and (3) direct feature transformation in the compressed domain. Experiments demonstrate that end-to-end training time reduces from days to hours, while memory utilization improves significantly, I/O overhead decreases, and instruction-level parallelism is enhanced.
This work addresses the lack of theoretical formalization and systematic implementation for “lossless” model compression. We propose LLC, the first general theoretical framework for provably lossless compression. Methodologically: (i) leveraging total differentials, we rigorously bound compression error and formally define the *lossless compression neighborhood* and higher-order analytical error bounds; (ii) we formulate quantization as a grouped knapsack problem, jointly optimizing layer-wise low-rank structures and quantization bit-widths to automatically determine the optimal lossless compression configuration. Rigorous evaluation across diverse architectures (ViT, ResNet) and datasets (ImageNet, CIFAR) confirms zero accuracy degradation post-compression, 1.8–3.2× inference speedup, and 40–65% memory reduction—without fine-tuning or heuristic design. Our core contribution is a theoretically grounded, computationally tractable, and broadly generalizable framework for lossless model compression.
To address the escalating network transmission and storage overheads in large language model (LLM) deployment, this paper introduces ZipNN—the first domain-specific, lossless compression framework tailored for neural network weights, supporting fully reversible compression and hardware-aware high-speed decompression. Unlike general-purpose compressors, ZipNN systematically exploits the unique statistical properties of neural weights for lossless compression, innovatively integrating entropy coding, fine-grained weight distribution modeling, block-wise adaptive quantization, and decoder scheduling optimization to yield architecture-adaptive compression variants. Evaluated on mainstream LLMs including Llama 3, ZipNN achieves over 17% greater space savings and 62% faster compression/decompression throughput compared to state-of-the-art general-purpose compressors (e.g., zstd). For a Hugging Face–scale platform, this translates to over 1 exabyte (EB) of monthly network traffic reduction.
This work addresses the growing complexity and lack of interpretability in deep image compression autoencoder models, which hinder the design of efficient architectures. For the first time, it systematically employs Jacobian analysis to examine the internal transformations of unbiased autoencoders, uncovering consistent and interpretable operational patterns that are prevalent across high-dimensional compression models. Building on these insights, the study identifies multiple semantically meaningful internal operations shared across diverse models and demonstrates their utility in constructing lightweight architectures that simultaneously achieve high compression performance and low computational complexity. This approach establishes a new paradigm for designing interpretable and efficient compression models grounded in analytically derived internal mechanisms.
This study addresses the storage and I/O bottlenecks faced by high-fidelity neural surrogate models, which stem from their reliance on large-scale training data, and tackles the challenge of quantifying the impact of lossy compression errors on model performance. The work proposes a novel method that leverages the inherent stochasticity of neural network training to assess error tolerance through uncertainty quantification, enabling a controllable trade-off between compression ratio and training efficiency while preserving model accuracy. Evaluated on two scientific simulation tasks, the approach achieves up to 39× data compression with up to 3× reduction in training time, while incurring negligible degradation in model quality.
This study addresses the optimization of general-purpose lossless compression under realistic resource constraints—specifically, memory usage capped at 8 GB and decompressor size limited to 1 MB—by organizing an international challenge based on a public training set and a hidden test set comprising 16 heterogeneous files. Performance is evaluated multidimensionally using compression ratio, compression/decompression time, Weissman score, and Pareto front analysis. The generalization capability of submissions is further assessed on external large-scale datasets, while Normalized Compression Distance (NCD) is employed to analyze inter-solution relationships. Among 117 valid submissions, several outperformed mainstream tools on external data, underscoring the critical role of advanced probabilistic modeling and effective multi-objective trade-offs in enhancing compression performance.
General-purpose compression algorithms often struggle to simultaneously achieve high compression ratios and high throughput with low overhead, whereas specialized compressors, while offering superior performance, incur high development and maintenance costs and suffer from limited applicability. This work proposes a novel “graph-based” compression framework that, for the first time, models the compression process as a modular composition of encoders and decoders represented by a directed acyclic graph. By integrating a self-describing format with a universal decoder (OpenZL), the framework unifies the generality of generic methods with the performance of specialized ones. The approach substantially reduces the cost of developing and deploying domain-specific compressors, outperforming mainstream general-purpose compressors in both compression ratio and speed across multiple real-world datasets. It remains competitive with deep learning–based methods while operating orders of magnitude faster, and internal adoption at Meta has reduced development cycles from months to days.
This work addresses the challenge of deploying deep neural networks on edge and embedded devices, where limited memory and computational resources necessitate a careful balance between model compression and performance. The authors propose a two-stage compression framework: first, joint pruning and quantization drastically reduce model size; second, a Mixture-of-Experts (MoE) mechanism dynamically routes inputs among multiple lightweight submodels to recover accuracy loss while preserving efficient inference. Notably, this study presents the first unified integration of pruning, quantization, and MoE architecture for effective ensemble-based compression. Experimental results demonstrate that the proposed method substantially reduces both parameter count and FLOPs of CNNs across multiple benchmark datasets, with only negligible degradation in accuracy.