🤖 AI Summary
This study addresses the coarse granularity and limited flexibility of fixed-bitwidth formats in neural network weight compression by proposing a calibration-free, entropy-coded compression framework that supports arbitrary target bitrates. The method integrates row-normalized E8 lattice quantization with a conditional probability model, enabling rapid lossless resolution selection and independently decodable tile encoding via sampling-based estimation. Combined with efficient GPU decoding, it achieves operator fusion for symbol decoding and numerical reconstruction while maintaining compatibility with BF16, FP8, and INT8 containers. Experimental results demonstrate that the proposed approach significantly outperforms NF4 on Z-Image-Turbo, reducing relative L2 error by approximately 24% at 4 bits per parameter, thereby substantially decreasing storage footprint while preserving low inference latency.
📝 Abstract
Weight compression helps large neural networks fit deployment memory budgets, but common fixed-width formats offer only coarse storage choices. Entropy coding supports finer rates, yet the achieved size depends on the quantized weight distribution and coding overhead. Exploiting this flexibility requires accurate rate selection and efficient weight reconstruction for inference. We present EntroPack, an entropy-coded weight compressor that supports arbitrary target bitrates without activation calibration or fine-tuning. It combines row-normalized $E_8$ lattice quantization with a conditional probability model of lattice coordinates. Sampled storage estimates select the quantization resolution without repeated full-stream encoding. The final coordinates are entropy-coded in independently decodable tiles, enabling fast, fused symbol decoding and numerical weight reconstruction on the GPU. EntroPack supports floating-point and integer weight containers, such as BF16, FP16, FP8, and INT8, with storage bitrate controlled independently of numerical precision. Online decoding adds latency that grows with weight count, making the method well suited to compute-intensive workloads such as diffusion denoising and Transformer prefill. Experiments demonstrate fast encoding and modest inference overhead in these settings. When compressing the linear-layer weights of the image generator Z-Image-Turbo, EntroPack achieves substantially lower weight and denoiser output errors than fixed-width formats at comparable storage rates, with modest denoising-step overhead. Targeting 4 bits per parameter, it achieves lower weight and denoiser output errors than NF4, including about 24% lower relative $L_2$ weight error, with less storage. Source code is available at https://github.com/modelscope/entropack.