Score
Designs and implements encoding architectures that map overlapping local patches into per-patch latent variables (often via patch-based VAEs), producing patch-local latent encodings that can be merged or stitched to reconstruct the whole input. These systems are built to preserve fine-grained local detail, maintain features in unedited regions, and enable localized analysis, reconstruction, or editing by combining patch latents.
Variational autoencoders (VAEs) often suffer from posterior collapse, degrading generative diversity; existing mitigation strategies rely on regularization trade-offs or architectural constraints, limiting generalizability and controllability. This paper proposes an architecture-agnostic method for localized posterior collapse control: we define a local collapse metric and introduce a latent reconstruction loss (LRL), leveraging the mathematical properties of injection and composition functions to enable end-to-end optimization within the variational inference framework. LRL requires no architectural modifications and jointly preserves reconstruction fidelity and latent identifiability. Experiments on MNIST, FashionMNIST, Omniglot, CelebA, and FFHQ demonstrate that our approach significantly alleviates posterior collapse, markedly improving sample diversity and distribution coverage. The method establishes a more robust and generalizable control paradigm for VAE training.
Existing representation-based encoder generative paradigms face two key challenges: (1) discriminative feature spaces lack compact regularization, causing diffusion sampling to deviate from the data manifold and yield structural distortions; and (2) encoders exhibit weak pixel-level reconstruction capability, limiting geometric and textural fidelity. To address these, we propose a semantic-pixel joint reconstruction objective, achieving—within a compact 16×16, 96-dimensional latent space—the first unified high-semantic and high-fidelity pixel reconstruction. Our method integrates dual reconstruction losses, a compact latent-space design, a unified representation-encoder-based T2I and editing diffusion architecture, and a VAE feature-space adaptation mechanism. Experiments demonstrate significant improvements in reconstruction quality, text-to-image generation, and image editing—achieving state-of-the-art performance—along with accelerated convergence. This validates the feasibility of efficiently transferring understanding-oriented encoders into robust, generative latent spaces.
In high-dimensional latent spaces, VAEs suffer from decoding failure when sampling from standard isotropic Gaussian priors due to the extreme sparsity of uniform distributions in high dimensions—a manifestation of the curse of dimensionality. Method: We propose Spherical Coordinate Reparameterization (SC-VAE), which explicitly maps latent variables onto the unit hypersphere and a radial dimension via a differentiable coordinate transformation. This enforces compact clustering of latent representations on the hypersphere while preserving end-to-end trainability—without architectural modifications or increased computational overhead. Contribution/Results: SC-VAE yields substantial improvements in generative performance for latent dimensions ≥128: FID improves by 30–50%, sample diversity increases, and the probability of generating valid samples under random prior sampling rises significantly. The method provides a lightweight, general-purpose, and theoretically grounded solution to degenerate generation in high-dimensional VAEs.
This work addresses the instability and degraded generation performance in video variational autoencoders when used with latent diffusion models, which arises from an excessive number of latent channels. To mitigate this issue, the authors propose a frequency-aware latent space compression method that selectively attenuates high-frequency components in the video latent representations, replacing conventional channel pruning. This approach preserves essential structural information while achieving the same compression ratio. The proposed method substantially improves reconstruction fidelity, enhances training stability of the diffusion model, and outperforms strong baseline methods in generation quality, thereby demonstrating the efficacy and advantages of frequency-guided compression for video generation tasks.
This work investigates the roles of masked-patch self-attention and masked-to-visible cross-attention in the MAE decoder for representation learning, revealing that image reconstruction primarily relies on global semantic representations extracted by the encoder—not on intra-masked-patch interactions within the decoder. Motivated by this finding, we propose CrossMAE: a streamlined framework that retains only the cross-attention mechanism while entirely removing self-attention among masked tokens in the decoder. This design is the first to empirically demonstrate that MAE’s effectiveness stems from the encoder’s strong global modeling capacity, challenging the prevailing assumption that decoder-side modeling of dependencies among masked patches is essential. Evaluated across ViT-S to ViT-H architectures, CrossMAE matches or surpasses standard MAE in performance while reducing GPU memory consumption by 37% and FLOPs by 42%. Code and pretrained models are publicly available.
This paper addresses the conceptual gap between variational autoencoders (VAEs) and vector-quantized VAEs (VQ-VAEs) in modeling continuous versus discrete latent representations. Methodologically, it proposes a novel autoencoder framework that eliminates both the KL divergence term and the reparameterization trick; instead, it explicitly enforces latent space compactness via learnable clustering centers and employs multi-vector outputs to jointly support continuous interpolation and discrete reconstruction. Key contributions include: (1) uncovering an intrinsic relationship between autoencoder generative fidelity and latent space compactness; (2) establishing a deterministic transition from VAEs to VQ-VAEs without relying on probabilistic assumptions; and (3) empirically validating smooth interpolation and stable reconstruction on MNIST, CelebA, and FashionMNIST. Experiments further reveal that naively increasing the number of output vectors leads to model degradation—manifesting as localized, patchwise discrete encoding—highlighting the critical role of architectural design.
This work challenges the conventional belief that variational autoencoder (VAE) encoders trained exclusively at low resolution (e.g., 256²) cannot generalize to high-resolution (e.g., 512²) image reconstruction. The authors identify and validate a counterintuitive phenomenon: compact student encoders obtained via knowledge distillation—despite being trained solely on low-resolution data—achieve superior reconstruction performance on unseen high-resolution inputs. By integrating input upsampling and output downsampling strategies, the proposed approach significantly improves key metrics including PSNR, SSIM, LPIPS, and rFID on ImageNet-256. This study demonstrates for the first time that distilled VAE models can generalize across resolutions without any high-resolution training, effectively inheriting the teacher model’s high-resolution representational capabilities and thereby redefining established assumptions about out-of-distribution generalization in generative modeling.
Standard variational autoencoders employ Gaussian priors, which struggle to align with data manifolds exhibiting non-Euclidean topologies—such as periodicity or boundedness—leading to distorted representations. This work proposes a topology-aware latent space modeling framework that constructs factorized prior distributions tailored to manifolds decomposable into products of circles, intervals, and lines, along with their finite group quotients. This design enables disentangled latent representations and analytically tractable KL divergences. By integrating differentiable coordinate transformations, group-invariant decoding, and anchor-point constraints, the approach ensures smooth gradients and topological consistency. To our knowledge, this is the first method to systematically align latent variable distributions with the intrinsic topology of data manifolds, supporting reparameterizable encoder–prior pairs and significantly outperforming Gaussian-prior baselines on synthetic manifolds as well as rotation- and cyclic-translation variants of MNIST.
This work addresses the limited reconstruction fidelity of existing latent diffusion models, which stems from the absence of low-level visual information—such as color and texture—in their semantic representations. To overcome this limitation, we propose LV-RAE (Low-level Visual Representation AutoEncoder), the first framework to effectively integrate fine-grained visual details into high-level semantic latents while preserving semantic structure. Our approach leverages a vision foundation model as the encoder and enhances decoder robustness through decoder fine-tuning, controlled noise injection, and latent space smoothing, thereby mitigating artifacts caused by perturbations in the latent variables. Experimental results demonstrate that LV-RAE significantly improves the perceptual quality and detail fidelity of generated images without compromising the model’s capacity for semantic abstraction.
This study challenges the presumed necessity of hierarchical quantization in Vector Quantized Variational Autoencoders (VQ-VAEs) for achieving high reconstruction quality. By systematically comparing single-layer and two-layer VQ-VAE architectures with matched representational capacity on high-resolution ImageNet, the work evaluates the actual contribution of hierarchical structure to reconstruction fidelity. Lightweight strategies—including data-driven codebook initialization, periodic resetting of inactive codebook vectors, and careful hyperparameter tuning—are employed to mitigate codebook collapse and enhance codebook utilization. Under controlled representational budgets and effective collapse suppression, the results demonstrate that a single-layer VQ-VAE can achieve reconstruction performance comparable to its hierarchical counterpart, thereby questioning the widely held assumption that hierarchical architectures are inherently superior.