Rethinking Generative Image Compression at Extremely Low Bitrates

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the semantic collapse problem in generative image compression at extremely low bitrates, where preserving recognizable content remains challenging. We identify that this issue fundamentally stems from loss conflicts and inefficient variational autoencoder (VAE) representations. To overcome these limitations, we propose the RAE-CoD framework, which employs a Representation Autoencoder (RAE) to directly align source-domain and compressed-domain representations, thereby preventing semantic degradation. Furthermore, a Compression-oriented Diffusion model (CoD) is constructed within the latent space to enable efficient decoding. Experimental results demonstrate that the proposed framework significantly reduces feature errors and Fréchet distances at an ultra-low bitrate of 16 bits, effectively preserving semantic integrity while facilitating a smooth transition from conditional to unconditional generation.
📝 Abstract
Generative image compression produces visually plausible reconstructions at low bitrates, yet their behavior as the rate approaches zero remains largely unexplored. When pushed below normal operating rates, representative codecs undergo semantic collapse: rather than gracefully losing source-specific detail, they produce malformed or unrecognizable content. Our analysis identifies two factors. As the bitrate decreases, reconstruction losses increasingly conflict with semantic objectives on gradients and visual results, while pixel-space and reconstruction-oriented VAE diffusion models become less efficient on semantic preservation. Guided by these findings, we introduce RAE-CoD, a compression-oriented diffusion (CoD) built in a representation autoencoder (RAE) space with direct alignment between compressed and source representations, preserving recognizable, naturally structured content for a $256\times256$ image with as few as 16 bits. We evaluate this framework using five vision foundation models (VFM) and a blinded vision-language model protocol. On MSCOCO-30K, RAE-CoD stands out from all evaluation. At 0.001-0.008 bpp, it reduces relative VFM feature MSE and Fréchet Distance ratio by at least 25.7% and 69.1% over the best competitors. Meanwhile, semantic recognizability and quality of the reconstructions remain nearly constant while source consistency falls smoothly, replacing abrupt semantic collapse with a graceful transition toward unconditional generation. Code will be released at https://github.com/LuizScarlet/RAE-CoD.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Image Compression
Extremely Low Bitrates
Representation Autoencoder
Compression-oriented Diffusion
Semantic Collapse
🔎 Similar Papers
No similar papers found.