🤖 AI Summary
This study addresses the loss of fine-grained details in pretrained visual representations during image generation and the difficulty of latent space modeling caused by multi-layer fusion. To tackle these challenges, we propose a Hierarchical Representation Autoencoder that introduces a residual budget mechanism to learn adaptive fusion across encoder layers. By constraining shallow-layer residual corrections via an upper bound on group norms, the method achieves efficient full-level fusion without complex hyperparameter tuning. Furthermore, integrating deep residual networks with hierarchical feature aggregation significantly enhances reconstruction fidelity while maintaining compatibility with mainstream generative frameworks. Experimental results demonstrate that our approach reduces the FID to 0.209 on ImageNet-256 and achieves a GenEval score of 87.70, substantially outperforming existing baseline models.
📝 Abstract
Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribution that is difficult to model. Existing fusion methods require empirical tuning of layer selection or staged optimization of fusion and decoding, increasing configuration effort or training complexity. We introduce HiRAE (Hierarchical Representation Autoencoder), which learns a hierarchical fusion framework over the full encoder hierarchy to improve reconstruction fidelity while maintaining compatibility with generative modeling. HiRAE groups encoder layers by depth and learns residual corrections to the deepest representation. Group-wise norm caps bound these corrections relative to the deep anchor, with tighter budgets for shallower groups. Our HiRAE-24 preserves the latent token count and channel dimension. On ImageNet-256, HiRAE-24 reduces reconstruction FID from 0.299 to 0.209 relative to RAEv2 while maintaining competitive guided generation quality. For text-to-image generation, HiRAE-24 improves alignment over RAEv2 on GenEval, DPG-Bench, and GenAI-Bench both before and after supervised fine-tuning. Under the same generator-training and evaluation protocol, post-fine-tuning GenEval increases from 84.86 to 87.70.