NeuralZip: Reusable Setup for Fast Lossless Compression

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational overhead in lossless model compression caused by redundant statistical modeling and encoding construction. We propose a fast, exact compression framework that leverages precomputed, reusable exponential distribution statistics. Methodologically, this work introduces the first cross-architecture transferable compression configuration, integrating exponential distribution grouping, shared Huffman coding, and packed exponent representation to substantially reduce computational redundancy. Experimental results demonstrate that the proposed approach achieves bit-exact reconstruction while accelerating compression speed by 1.8× to 21× and reducing GPU memory consumption by 27.5%. These findings establish a new paradigm for efficient model deployment.
📝 Abstract
Lossless compression can reduce the storage and movement of model weights without changing their floating-point values, but repeated statistical analysis and code construction add computational overhead. We study whether the statistical structure of exponents can be prepared once and reused. For this, we introduce NeuralZip, which groups chunks with similar exponent distributions, shares Huffman codes, and selectively represents recurring exponent tuples using packed exponents, thereby achieving additional moderate compression ratios. A setup chooses these representations before subsequent encodings, while every encoding still processes the current tensor values. In floating-point model checkpoints, post-setup compression is 1.81-21.33$\times$ faster than the baselines and achieves exact bit-to-bit reconstruction. We show that this setup can be precomputed and transferred from another compatible architecture, preserving similar compression ratios and avoiding the need to amortize setup costs. Therefore, compression adaptation is transferable and reusable. Training checkpoints demonstrate continued reuse as the weights evolve. Finally, GPU experiments reduce active memory usage by up to 27.5$\%$ while reproducing the logits exactly.
Problem

Research questions and friction points this paper is trying to address.

Lossless compression
Computational overhead
Reusable setup
Model weights
Floating-point
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lossless Compression
NeuralZip
Reusable Setup
Huffman Coding
Memory Efficiency
💼 Related Jobs
No related jobs found.
M
Martín Bravo
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE
S
Samuel Horváth
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE
Gonzalo Navarro
Gonzalo Navarro
University of Chile
algorithms and data structurestext searchingcompressiongraph databasessimilarity search
Andrés Abeliuk
Andrés Abeliuk
Department of Computer Science, University of Chile
Machine LearningComputational Social ScienceNetwork ScienceAI & Society