🤖 AI Summary
Neural image compression faces challenges in practical deployment due to high decoder computational complexity. To address this, we propose a lightweight compression-reconstruction framework integrating low-rank representation with vector-quantized variational autoencoders (VQ-VAEs). Our core innovation replaces conventional high-dimensional convolutional upsampling in the VQ-VAE decoder with a learnable low-rank projection module and designs a streamlined decoder architecture, substantially reducing parameter count and FLOPs during decoding. Experiments demonstrate that our method achieves PSNR and MS-SSIM comparable to state-of-the-art approaches while accelerating decoding by 2.3–4.1× and reducing GPU memory consumption by ~60%. To the best of our knowledge, this is the first work to systematically embed low-rank modeling into the VQ-VAE decoding pipeline, effectively balancing high-fidelity reconstruction with ultra-efficient decoding.
📝 Abstract
Image compression and reconstruction are crucial for various digital applications. While contemporary neural compression methods achieve impressive compression rates, the adoption of such technology has been largely hindered by the complexity and large computational costs of the convolution-based decoders during data reconstruction. To address the decoder bottleneck in neural compression, we develop a new compression-reconstruction framework based on incorporating low-rank representation in an autoencoder with vector quantization. We demonstrated that performing a series of computationally efficient low-rank operations on the learned latent representation of images can efficiently reconstruct the data with high quality. Our approach dramatically reduces the computational overhead in the decoding phase of neural compression/reconstruction, essentially eliminating the decoder compute bottleneck while maintaining high fidelity of image outputs.