🤖 AI Summary
This study addresses the high computational cost and low throughput of autoregressive decoding in large language model (LLM)-based data compression. To this end, we propose an efficient lossless compression framework that integrates diffusion LLMs with multi-token prediction. The core innovation lies in introducing a hypernetwork architecture that dynamically generates data-specific parameter updates, enabling personalized adaptation without fine-tuning and effectively mitigating the trade-off between high throughput and low compression ratio. Experimental results demonstrate that the proposed method significantly outperforms existing baselines in both compression ratio and inference speed, achieving highly efficient lossless data compression.
📝 Abstract
Large language models (LLMs) have shown strong potential for lossless data compression, but existing approaches are constrained by the high computational cost and low throughput of autoregressive decoding. We propose HyperZip, an efficient and scalable LLM-based compression framework that leverages diffusion-based LLMs (dLLMs) with Multi-Token Prediction (MTP) to accelerate LLM-based data compression processes. We identify a trade-off in diffusion-based compression, where increasing decoding throughput degrades the compression rate. To mitigate this trade-off, HyperZip employs a hypernetwork to generate data-specific updates from a context representation, adapting the dLLM to the target data without costly fine-tuning, resulting in a low compression rate and high throughput. Extensive experiments show that HyperZip achieves a superior trade-off between compression rate and speed compared with state-of-the-art baselines.