Ultra-Efficient Decoding for End-to-End Neural Compression and Reconstruction

📅 2025-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Neural image compression faces challenges in practical deployment due to high decoder computational complexity. To address this, we propose a lightweight compression-reconstruction framework integrating low-rank representation with vector-quantized variational autoencoders (VQ-VAEs). Our core innovation replaces conventional high-dimensional convolutional upsampling in the VQ-VAE decoder with a learnable low-rank projection module and designs a streamlined decoder architecture, substantially reducing parameter count and FLOPs during decoding. Experiments demonstrate that our method achieves PSNR and MS-SSIM comparable to state-of-the-art approaches while accelerating decoding by 2.3–4.1× and reducing GPU memory consumption by ~60%. To the best of our knowledge, this is the first work to systematically embed low-rank modeling into the VQ-VAE decoding pipeline, effectively balancing high-fidelity reconstruction with ultra-efficient decoding.

Technology Category

Computer Vision: Representation Learning for VisionMachine Learning: Deep Generative Models & AutoencodersCognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingGraph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphs
📝 Abstract
Image compression and reconstruction are crucial for various digital applications. While contemporary neural compression methods achieve impressive compression rates, the adoption of such technology has been largely hindered by the complexity and large computational costs of the convolution-based decoders during data reconstruction. To address the decoder bottleneck in neural compression, we develop a new compression-reconstruction framework based on incorporating low-rank representation in an autoencoder with vector quantization. We demonstrated that performing a series of computationally efficient low-rank operations on the learned latent representation of images can efficiently reconstruct the data with high quality. Our approach dramatically reduces the computational overhead in the decoding phase of neural compression/reconstruction, essentially eliminating the decoder compute bottleneck while maintaining high fidelity of image outputs.
Problem

Research questions and friction points this paper is trying to address.

Addressing decoder complexity bottleneck in neural image compression
Reducing computational costs of convolution-based reconstruction methods
Maintaining high fidelity while eliminating decoder compute bottleneck
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses low-rank representation in autoencoder with vector quantization
Performs efficient low-rank operations on learned latent representations
Eliminates decoder compute bottleneck while maintaining high fidelity
🔎 Similar Papers
No similar papers found.
E
Ethan G. Rogers
Department of Computer Enginering, Iowa State University, Ames, IA 50011
C
Cheng Wang
Department of Computer Enginering, Iowa State University, Ames, IA 50011