🤖 AI Summary
This study addresses the limitation that quantization and generative residuals impose on reconstruction quality in autoregressive image coding by proposing the ResARC framework. This method introduces a novel explicit dual residual compensation mechanism, leveraging diffusion models to compensate for quantization residuals while compressing and transmitting generative residuals. By integrating diffusion transformers with learned codecs, ResARC achieves high-fidelity context-based reconstruction without requiring additional side information. Experimental results demonstrate that ResARC significantly enhances distributional fidelity at ultra-low bitrates while maintaining highly competitive perceptual similarity.
📝 Abstract
Progressive autoregressive image codecs provide an appealing paradigm for generative compression by quantizing continuous latents into discrete tokens, transmitting coarse-to-fine prefix tokens and generating the remaining suffix tokens at the decoder. However, their reconstruction quality is fundamentally limited by two residuals introduced along this pipeline: the quantization residual, arising from information loss during discrete tokenization, and the generation residual, resulting from imperfect autoregressive generation of the suffix tokens. To address these limitations, we introduce ResARC, a residual-aware autoregressive codec that explicitly compensates for both residuals at the decoder. Specifically, we generate the quantization residual with a diffusion transformer conditioned on the autoregressive decoding context, while requiring no additional side information. In parallel, we compute the generation residual at the encoder and employ a learned Generation Residual Codec to efficiently compress and transmit it for decoder-side correction. The recovered residuals are then integrated with the reconstructed latent representation and decoded through an adapted VAE decoder. Extensive experiments demonstrate that ResARC achieves competitive perceptual similarity while substantially improving distributional fidelity over leading generative codecs across the ultra-low bitrate regime. Code and models will be released soon.