🤖 AI Summary
This study addresses the high computational overhead in lossless model compression caused by redundant statistical modeling and encoding construction. We propose a fast, exact compression framework that leverages precomputed, reusable exponential distribution statistics. Methodologically, this work introduces the first cross-architecture transferable compression configuration, integrating exponential distribution grouping, shared Huffman coding, and packed exponent representation to substantially reduce computational redundancy. Experimental results demonstrate that the proposed approach achieves bit-exact reconstruction while accelerating compression speed by 1.8× to 21× and reducing GPU memory consumption by 27.5%. These findings establish a new paradigm for efficient model deployment.
📝 Abstract
Lossless compression can reduce the storage and movement of model weights without changing their floating-point values, but repeated statistical analysis and code construction add computational overhead. We study whether the statistical structure of exponents can be prepared once and reused. For this, we introduce NeuralZip, which groups chunks with similar exponent distributions, shares Huffman codes, and selectively represents recurring exponent tuples using packed exponents, thereby achieving additional moderate compression ratios. A setup chooses these representations before subsequent encodings, while every encoding still processes the current tensor values. In floating-point model checkpoints, post-setup compression is 1.81-21.33$\times$ faster than the baselines and achieves exact bit-to-bit reconstruction. We show that this setup can be precomputed and transferred from another compatible architecture, preserving similar compression ratios and avoiding the need to amortize setup costs. Therefore, compression adaptation is transferable and reusable. Training checkpoints demonstrate continued reuse as the weights evolve. Finally, GPU experiments reduce active memory usage by up to 27.5$\%$ while reproducing the logits exactly.