compression-aware representation learning

Designs, builds, and evaluates algorithms and models that produce or exploit compact, compression-aware representations and compressed models while preserving information needed for downstream tasks. This includes designing lossy and lossless compression methods for embeddings and signals, compression-aware pretraining and contrastive objectives using compressed echoes, compact autoencoder architectures, hardware-aware model compression and optimization, sensitivity-profile guided compression, and metrics/procedures for compression evaluation.

compression-awarerepresentationlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Challenges and Solutions in Selecting Optimal Lossless Data Compression Algorithms

Sep 23, 2025
MA
Md. Atiqur Rahman
🏛️ East West University | Bangladesh University of Business and Technology (BUBT)

Lossless compression algorithms face inherent trade-offs among multidimensional performance metrics—particularly compression ratio and encoding/decoding speed—posing challenges in latency-sensitive, high-fidelity applications such as medical imaging. Method: This paper introduces the first unified, quantifiable multi-objective evaluation framework for lossless compression, employing normalized weighted modeling to dynamically balance compression ratio and speed across diverse data modalities (e.g., images, text). Contribution/Results: Its key innovation lies in a standardized multi-objective scoring model that bridges the gap between theoretical metrics and practical deployment requirements. Extensive experiments demonstrate the framework’s robustness in identifying scenario-optimal compressors: learned codecs achieve superior compression ratios, while traditional algorithms retain advantages in speed-critical tasks. The framework enables principled, application-aware algorithm selection without requiring domain-specific re-engineering.

Balancing compression ratio, encoding speed, and decoding speed requirementsProviding objective comparisons for diverse applications like medical imagingSelecting optimal lossless compression algorithms with conflicting performance trade-offs

Morphing-based Compression for Data-centric ML Pipelines

Apr 15, 2025
SB
Sebastian Baunsgaard
🏛️ Technische Universität Berlin

Existing lossless matrix compression methods fail to effectively capture structural redundancies introduced during data cleaning, augmentation, and feature engineering, leading to suboptimal efficiency in data-centric ML pipelines. This paper introduces BWARE, the first framework to deeply embed lossless compression within the outer loop of data engineering—enabling end-to-end co-design of compression and feature transformation. Its key contributions are: (1) a workload-aware compression mechanism supporting lightweight, on-the-fly morphing without decompression; (2) a column correlation- and sparsity-aware morphing mapping; and (3) direct feature transformation in the compressed domain. Experiments demonstrate that end-to-end training time reduces from days to hours, while memory utilization improves significantly, I/O overhead decreases, and instruction-level parallelism is enhanced.

Efficiently compressing data-centric ML pipeline matricesLeveraging structural transformations for lossless compressionReducing ML pipeline runtime via workload-optimized compression

This work addresses the lack of theoretical formalization and systematic implementation for “lossless” model compression. We propose LLC, the first general theoretical framework for provably lossless compression. Methodologically: (i) leveraging total differentials, we rigorously bound compression error and formally define the *lossless compression neighborhood* and higher-order analytical error bounds; (ii) we formulate quantization as a grouped knapsack problem, jointly optimizing layer-wise low-rank structures and quantization bit-widths to automatically determine the optimal lossless compression configuration. Rigorous evaluation across diverse architectures (ViT, ResNet) and datasets (ImageNet, CIFAR) confirms zero accuracy degradation post-compression, 1.8–3.2× inference speedup, and 40–65% memory reduction—without fine-tuning or heuristic design. Our core contribution is a theoretically grounded, computationally tractable, and broadly generalizable framework for lossless model compression.

Defines error boundaries for lossless compression to minimize model degradation.Stabilizes lossless model compression to reduce complexity without performance loss.Systematically determines compression impact on model performance using a theoretical framework.

ZipNN: Lossless Compression for AI Models

Nov 07, 2024
MH
Moshik Hershcovitch
🏛️ IBM | Tel Aviv University | Boston University | MIT | Dartmouth College

To address the escalating network transmission and storage overheads in large language model (LLM) deployment, this paper introduces ZipNN—the first domain-specific, lossless compression framework tailored for neural network weights, supporting fully reversible compression and hardware-aware high-speed decompression. Unlike general-purpose compressors, ZipNN systematically exploits the unique statistical properties of neural weights for lossless compression, innovatively integrating entropy coding, fine-grained weight distribution modeling, block-wise adaptive quantization, and decoder scheduling optimization to yield architecture-adaptive compression variants. Evaluated on mainstream LLMs including Llama 3, ZipNN achieves over 17% greater space savings and 62% faster compression/decompression throughput compared to state-of-the-art general-purpose compressors (e.g., zstd). For a Hugging Face–scale platform, this translates to over 1 exabyte (EB) of monthly network traffic reduction.

Achieving significant space savings and faster compression/decompression speedsLossless compression for reducing AI model storage and network burdenSpecialized compression variants to enhance model compressibility effectiveness

Latest Papers

What's happening recently
View more

This work addresses the growing complexity and lack of interpretability in deep image compression autoencoder models, which hinder the design of efficient architectures. For the first time, it systematically employs Jacobian analysis to examine the internal transformations of unbiased autoencoders, uncovering consistent and interpretable operational patterns that are prevalent across high-dimensional compression models. Building on these insights, the study identifies multiple semantically meaningful internal operations shared across diverse models and demonstrates their utility in constructing lightweight architectures that simultaneously achieve high compression performance and low computational complexity. This approach establishes a new paradigm for designing interpretable and efficient compression models grounded in analytically derived internal mechanisms.

autoencodersimage compressioninterpretability

This study addresses the storage and I/O bottlenecks faced by high-fidelity neural surrogate models, which stem from their reliance on large-scale training data, and tackles the challenge of quantifying the impact of lossy compression errors on model performance. The work proposes a novel method that leverages the inherent stochasticity of neural network training to assess error tolerance through uncertainty quantification, enabling a controllable trade-off between compression ratio and training efficiency while preserving model accuracy. Evaluated on two scientific simulation tasks, the approach achieves up to 39× data compression with up to 3× reduction in training time, while incurring negligible degradation in model quality.

compression errorgenerative surrogate modelinglossy compression

This study addresses the optimization of general-purpose lossless compression under realistic resource constraints—specifically, memory usage capped at 8 GB and decompressor size limited to 1 MB—by organizing an international challenge based on a public training set and a hidden test set comprising 16 heterogeneous files. Performance is evaluated multidimensionally using compression ratio, compression/decompression time, Weissman score, and Pareto front analysis. The generalization capability of submissions is further assessed on external large-scale datasets, while Normalized Compression Distance (NCD) is employed to analyze inter-solution relationships. Among 117 valid submissions, several outperformed mainstream tools on external data, underscoring the critical role of advanced probabilistic modeling and effective multi-objective trade-offs in enhancing compression performance.

Algorithmic Information Theorycompression benchmarkgeneralization

General-purpose compression algorithms often struggle to simultaneously achieve high compression ratios and high throughput with low overhead, whereas specialized compressors, while offering superior performance, incur high development and maintenance costs and suffer from limited applicability. This work proposes a novel “graph-based” compression framework that, for the first time, models the compression process as a modular composition of encoders and decoders represented by a directed acyclic graph. By integrating a self-describing format with a universal decoder (OpenZL), the framework unifies the generality of generic methods with the performance of specialized ones. The approach substantially reduces the cost of developing and deploying domain-specific compressors, outperforming mainstream general-purpose compressors in both compression ratio and speed across multiple real-world datasets. It remains competitive with deep learning–based methods while operating orders of magnitude faster, and internal adoption at Meta has reduced development cycles from months to days.

application-specific compressorslossless compressionmaintainability

This work addresses the challenge of deploying deep neural networks on edge and embedded devices, where limited memory and computational resources necessitate a careful balance between model compression and performance. The authors propose a two-stage compression framework: first, joint pruning and quantization drastically reduce model size; second, a Mixture-of-Experts (MoE) mechanism dynamically routes inputs among multiple lightweight submodels to recover accuracy loss while preserving efficient inference. Notably, this study presents the first unified integration of pruning, quantization, and MoE architecture for effective ensemble-based compression. Experimental results demonstrate that the proposed method substantially reduces both parameter count and FLOPs of CNNs across multiple benchmark datasets, with only negligible degradation in accuracy.

computational resourcesedge devicesmemory constraints

Hot Scholars

SD

Sheng Di

Argonne National Labratory, IEEE Senior Member
HPCData CompressionResilienceCloud/Grid Computing/P2P
WL

Weisi Lin

President's Chair Professor in Computer Science, CCDS, Nanyang Technological Unversity
Perception-inspired signal modelingperceptual multimedia quality evaluationvideo compressionimage processing & analysis
FC

Franck Cappello

Argonne National Laboratory, IEEE Fellow
Parallel ProcessingParallel ComputingHigh Performance ComputingFault Tolerance
JT

Jiannan Tian

Assistant Professor, Oakland University
HPC/AIlarge-scale data processing and analyticsHW-accelerated compression
HG

Hanqi Guo

The Ohio State University
Data visualization and analysis