multi-patch latent encoding

Designs and implements encoding architectures that map overlapping local patches into per-patch latent variables (often via patch-based VAEs), producing patch-local latent encodings that can be merged or stitched to reconstruct the whole input. These systems are built to preserve fine-grained local detail, maintain features in unedited regions, and enable localized analysis, reconstruction, or editing by combining patch latents.

multi-patchlatentencoding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Toward Architecture-Agnostic Local Control of Posterior Collapse in VAEs

Aug 17, 2025
HS
Hyunsoo Song
🏛️ National Institute for Mathematical Sciences | Seoul National University | KyungHee University

Variational autoencoders (VAEs) often suffer from posterior collapse, degrading generative diversity; existing mitigation strategies rely on regularization trade-offs or architectural constraints, limiting generalizability and controllability. This paper proposes an architecture-agnostic method for localized posterior collapse control: we define a local collapse metric and introduce a latent reconstruction loss (LRL), leveraging the mathematical properties of injection and composition functions to enable end-to-end optimization within the variational inference framework. LRL requires no architectural modifications and jointly preserves reconstruction fidelity and latent identifiability. Experiments on MNIST, FashionMNIST, Omniglot, CelebA, and FFHQ demonstrate that our approach significantly alleviates posterior collapse, markedly improving sample diversity and distribution coverage. The method establishes a more robust and generalizable control paradigm for VAE training.

Addressing posterior collapse in VAEs to enhance sample diversityOvercoming architectural constraints for latent identifiability in VAEsProposing architecture-agnostic loss to control posterior collapse

Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing

Dec 19, 2025
SZ
Shilong Zhang
🏛️ The University of Hong Kong | Adobe Research | University of Chinese Academy of Sciences

Existing representation-based encoder generative paradigms face two key challenges: (1) discriminative feature spaces lack compact regularization, causing diffusion sampling to deviate from the data manifold and yield structural distortions; and (2) encoders exhibit weak pixel-level reconstruction capability, limiting geometric and textural fidelity. To address these, we propose a semantic-pixel joint reconstruction objective, achieving—within a compact 16×16, 96-dimensional latent space—the first unified high-semantic and high-fidelity pixel reconstruction. Our method integrates dual reconstruction losses, a compact latent-space design, a unified representation-encoder-based T2I and editing diffusion architecture, and a VAE feature-space adaptation mechanism. Experiments demonstrate significant improvements in reconstruction quality, text-to-image generation, and image editing—achieving state-of-the-art performance—along with accelerated convergence. This validates the feasibility of efficiently transferring understanding-oriented encoders into robust, generative latent spaces.

Adapting discriminative encoders for generative tasks with semantic-pixel regularizationEnabling compact, semantically rich latents for text-to-image generation and editingOvercoming off-manifold latents and weak reconstruction in representation encoders

Improving the Generation of VAEs with High Dimensional Latent Spaces by the use of Hyperspherical Coordinates

Jul 21, 2025
AA
Alejandro Ascarate
🏛️ Queensland University of Technology | Data61 | CSIRO

In high-dimensional latent spaces, VAEs suffer from decoding failure when sampling from standard isotropic Gaussian priors due to the extreme sparsity of uniform distributions in high dimensions—a manifestation of the curse of dimensionality. Method: We propose Spherical Coordinate Reparameterization (SC-VAE), which explicitly maps latent variables onto the unit hypersphere and a radial dimension via a differentiable coordinate transformation. This enforces compact clustering of latent representations on the hypersphere while preserving end-to-end trainability—without architectural modifications or increased computational overhead. Contribution/Results: SC-VAE yields substantial improvements in generative performance for latent dimensions ≥128: FID improves by 30–50%, sample diversity increases, and the probability of generating valid samples under random prior sampling rises significantly. The method provides a lightweight, general-purpose, and theoretically grounded solution to degenerate generation in high-dimensional VAEs.

Addressing latent sparsity using hyperspherical coordinatesEnhancing meaningful data generation from random latent vectorsImproving VAE generation in high-dimensional latent spaces

This work addresses the instability and degraded generation performance in video variational autoencoders when used with latent diffusion models, which arises from an excessive number of latent channels. To mitigate this issue, the authors propose a frequency-aware latent space compression method that selectively attenuates high-frequency components in the video latent representations, replacing conventional channel pruning. This approach preserves essential structural information while achieving the same compression ratio. The proposed method substantially improves reconstruction fidelity, enhances training stability of the diffusion model, and outperforms strong baseline methods in generation quality, thereby demonstrating the efficacy and advantages of frequency-guided compression for video generation tasks.

generative performancelatent channelslatent diffusion models

Rethinking Patch Dependence for Masked Autoencoders

Jan 25, 2024
LF
Letian Fu
🏛️ UC Berkeley | UCSF

This work investigates the roles of masked-patch self-attention and masked-to-visible cross-attention in the MAE decoder for representation learning, revealing that image reconstruction primarily relies on global semantic representations extracted by the encoder—not on intra-masked-patch interactions within the decoder. Motivated by this finding, we propose CrossMAE: a streamlined framework that retains only the cross-attention mechanism while entirely removing self-attention among masked tokens in the decoder. This design is the first to empirically demonstrate that MAE’s effectiveness stems from the encoder’s strong global modeling capacity, challenging the prevailing assumption that decoder-side modeling of dependencies among masked patches is essential. Evaluated across ViT-S to ViT-H architectures, CrossMAE matches or surpasses standard MAE in performance while reducing GPU memory consumption by 37% and FLOPs by 42%. Code and pretrained models are publicly available.

Challenges mask token interaction necessity in masked pretrainingExamines inter-patch dependencies in MAE decoders for representation learningProposes CrossMAE using only cross-attention to reduce computation

Latest Papers

What's happening recently
View more

Rethinking VAE: From Continuous to Discrete Representations Without Probabilistic Assumptions

Jul 23, 2025
SS
Songxuan Shi
🏛️ Beijing University of Technology

This paper addresses the conceptual gap between variational autoencoders (VAEs) and vector-quantized VAEs (VQ-VAEs) in modeling continuous versus discrete latent representations. Methodologically, it proposes a novel autoencoder framework that eliminates both the KL divergence term and the reparameterization trick; instead, it explicitly enforces latent space compactness via learnable clustering centers and employs multi-vector outputs to jointly support continuous interpolation and discrete reconstruction. Key contributions include: (1) uncovering an intrinsic relationship between autoencoder generative fidelity and latent space compactness; (2) establishing a deterministic transition from VAEs to VQ-VAEs without relying on probabilistic assumptions; and (3) empirically validating smooth interpolation and stable reconstruction on MNIST, CelebA, and FashionMNIST. Experiments further reveal that naively increasing the number of output vectors leads to model degradation—manifesting as localized, patchwise discrete encoding—highlighting the critical role of architectural design.

Addressing blurriness in interpolations via compact latent spacesEnhancing AE generative ability via latent space clusteringExploring VAE-VQVAE connections without probabilistic assumptions

This work challenges the conventional belief that variational autoencoder (VAE) encoders trained exclusively at low resolution (e.g., 256²) cannot generalize to high-resolution (e.g., 512²) image reconstruction. The authors identify and validate a counterintuitive phenomenon: compact student encoders obtained via knowledge distillation—despite being trained solely on low-resolution data—achieve superior reconstruction performance on unseen high-resolution inputs. By integrating input upsampling and output downsampling strategies, the proposed approach significantly improves key metrics including PSNR, SSIM, LPIPS, and rFID on ImageNet-256. This study demonstrates for the first time that distilled VAE models can generalize across resolutions without any high-resolution training, effectively inheriting the teacher model’s high-resolution representational capabilities and thereby redefining established assumptions about out-of-distribution generalization in generative modeling.

Cross-resolution GeneralizationKnowledge DistillationLatent Manifold

Standard variational autoencoders employ Gaussian priors, which struggle to align with data manifolds exhibiting non-Euclidean topologies—such as periodicity or boundedness—leading to distorted representations. This work proposes a topology-aware latent space modeling framework that constructs factorized prior distributions tailored to manifolds decomposable into products of circles, intervals, and lines, along with their finite group quotients. This design enables disentangled latent representations and analytically tractable KL divergences. By integrating differentiable coordinate transformations, group-invariant decoding, and anchor-point constraints, the approach ensures smooth gradients and topological consistency. To our knowledge, this is the first method to systematically align latent variable distributions with the intrinsic topology of data manifolds, supporting reparameterizable encoder–prior pairs and significantly outperforming Gaussian-prior baselines on synthetic manifolds as well as rotation- and cyclic-translation variants of MNIST.

latent spacemanifold representationnon-Euclidean topology

This work addresses the limited reconstruction fidelity of existing latent diffusion models, which stems from the absence of low-level visual information—such as color and texture—in their semantic representations. To overcome this limitation, we propose LV-RAE (Low-level Visual Representation AutoEncoder), the first framework to effectively integrate fine-grained visual details into high-level semantic latents while preserving semantic structure. Our approach leverages a vision foundation model as the encoder and enhances decoder robustness through decoder fine-tuning, controlled noise injection, and latent space smoothing, thereby mitigating artifacts caused by perturbations in the latent variables. Experimental results demonstrate that LV-RAE significantly improves the perceptual quality and detail fidelity of generated images without compromising the model’s capacity for semantic abstraction.

decoder sensitivitylatent diffusion modelslow-level information

This study challenges the presumed necessity of hierarchical quantization in Vector Quantized Variational Autoencoders (VQ-VAEs) for achieving high reconstruction quality. By systematically comparing single-layer and two-layer VQ-VAE architectures with matched representational capacity on high-resolution ImageNet, the work evaluates the actual contribution of hierarchical structure to reconstruction fidelity. Lightweight strategies—including data-driven codebook initialization, periodic resetting of inactive codebook vectors, and careful hyperparameter tuning—are employed to mitigate codebook collapse and enhance codebook utilization. Under controlled representational budgets and effective collapse suppression, the results demonstrate that a single-layer VQ-VAE can achieve reconstruction performance comparable to its hierarchical counterpart, thereby questioning the widely held assumption that hierarchical architectures are inherently superior.

codebook collapsehierarchical quantizationreconstruction fidelity

Hot Scholars

DS

Dvir Samuel

PhD Student, Bar Ilan University
Machine LearningDeep LearningFew Shot LearningLong-tail Learning
HL

Hongdong Li

Professor of Computer Vision and Machine Learning, ANU, and Amazon IML
Computer VisionMachine LearningVR/ARArtificial Intelligence
YS

Yiren Song

PH.D student, National University of Singapore
Generative AIDiffusionUnified model