🤖 AI Summary
This study addresses the suboptimality arising from the decoupling of pruning and quantization in SVD-based compression by proposing a unified co-optimization framework. Methodologically, it introduces a differentiable component-wise bit-width learning mechanism that automatically assigns zero bits to low-importance components, thereby achieving precise pruning. By integrating SVD decomposition with a joint optimization algorithm, the framework enables end-to-end synergy between pruning and quantization. Experimental results demonstrate that under extreme 1.61-bit compression, the proposed approach significantly outperforms two-stage baseline methods. This work establishes a new paradigm for highly efficient compression of large models.
📝 Abstract
SVD-based pruning and quantization have recently emerged as a promising strategy for the ultra-efficient compression of large language models. In these methods, compression is performed in two stages: components are first truncated, and the remaining ones are subsequently quantized. Although this decoupled pipeline benefits from both pruning and quantization, it requires separate optimization for each stage and fails to fully exploit their balance, which can lead to suboptimal performance under aggressive compression. To address this limitation, we propose a new LLM compression method that co-optimizes pruning and quantization in a unified framework. Our key idea is a differentiable method for learning component-wise bit-widths, allowing less important components to be assigned 0-bit precision and pruned away. Notably, our method performs favorably against two-stage baselines, even when subjected to extreme quantization settings ($1.61$ bits) designed for ultra-efficiency. Code: https://github.com/MMAI-Laboratory/DBW.