masked latent prediction

Designs and implements models and training objectives that hide (mask) parts of inputs or intermediate representations and predict or reconstruct the missing latent tokens, feature embeddings, or scene-level representations from the visible context using generative reconstruction, embedding-alignment, or self-distillation losses. Measures and analyzes reconstruction fidelity and embedding alignment in latent space and the effects of these objectives on representational invariance and downstream robustness.

maskedlatentprediction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.53
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$207K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Latent Denoising Makes Good Visual Tokenizers

Jul 21, 2025
JY
Jiawei Yang
🏛️ USC | MIT CSAIL | Google DeepMind | OpenAI

Visual tokenizers are critical for generative modeling, yet their design lacks explicit alignment with the denoising reconstruction objective. This paper introduces a latent denoising-driven tokenizer design paradigm—unifying tokenizer training with the core generative task of reconstructing clean latent representations from noisy or masked inputs for the first time. To this end, we propose the Latent Denoising Tokenizer (l-DeTok), an autoencoder-based architecture jointly optimized with interpolation-based Gaussian noise injection and random masking reconstruction losses, thereby significantly enhancing token embeddability and robustness. Evaluated on ImageNet at 256×256 resolution, l-DeTok consistently outperforms standard tokenizers across six state-of-the-art generative models, yielding substantial improvements in generation quality. Our approach establishes a principled framework for tokenizer design grounded in the fundamental denoising objective of latent diffusion and masked autoencoding models.

Align tokenizer embeddings with denoising objectiveEnhance tokenizer performance for generative modelsImprove reconstruction of corrupted latent embeddings

Masked Image Modeling: A Survey

Aug 13, 2024
VH
Vlad Hondru
🏛️ University of Bucharest | Amazon | University of Trento

Existing research on masked image modeling (MIM) for self-supervised visual representation learning lacks a unified formalization of pretraining paradigms and standardized, comparable evaluation. Method: We formally categorize MIM into two principal paradigms—reconstruction-based and contrastive-based—and construct an interpretable, hierarchical taxonomy via expert curation and agglomerative clustering. We conduct systematic, unified benchmarking of over 20 state-of-the-art models on ImageNet and other standard datasets. Contribution/Results: Our analysis identifies critical open challenges—including cross-paradigm integration, long-tailed masking strategies, and compute-accuracy trade-offs. To foster reproducibility and standardization, we publicly release a structured literature repository and an extensible evaluation framework on GitHub. This work establishes foundational infrastructure for rigorous, comparable advancement in MIM research.

Automatic LearningComputer VisionMasked Image Modeling

Joint Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self Supervised Learning

May 18, 2025
HV
Hugues Van Assel
🏛️ Genentech | Meta AI | Brown University

The lack of principled criteria for choosing between contrastive joint-embedding and generative reconstruction paradigms in self-supervised learning (SSL) hinders theoretical understanding and practical design. Method: We conduct a rigorous theoretical analysis under linear model assumptions, deriving closed-form solutions for both paradigms and explicitly modeling the view-generation process to characterize the impact of data augmentations and nuisance features on representation learning. Results: Our analysis reveals that joint-embedding achieves asymptotically optimal performance under strong nuisance features with weaker alignment requirements, whereas reconstruction is inherently sensitive to such features. We further derive the minimal necessary condition linking augmentation strength and feature alignment. This work provides the first quantitative, interpretable explanation for the empirical superiority of joint-embedding over reconstruction on complex real-world data, establishing the first theoretically grounded, explainable guidance for SSL paradigm selection.

Analyze impact of view generation on representationsCompare reconstruction and joint embedding in SSLDetermine optimal SSL paradigm for irrelevant features

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Oct 09, 2024
SY
Sihyun Yu
🏛️ KAIST | Korea University | Scaled Foundations | New York University

Diffusion models suffer from inefficient representation learning and limited generation quality due to semantically impoverished latent spaces. To address this, we propose REPA (Representation Alignment), a novel regularization method that explicitly aligns denoising latent states—corrupted by noise—with clean-image representations extracted from high-quality external vision encoders (e.g., CLIP or DINO) within diffusion Transformers (DiT/SiT). This alignment is enforced via a projection-based loss, optimized end-to-end to enhance semantic consistency in the latent space. Experiments demonstrate that REPA accelerates SiT training by over 17.5×, enabling a SiT model to match the performance of a 7M-step SiT-XL within fewer than 400K steps. With classifier-free guidance (CFG), the method achieves an FID of 1.42—setting a new state-of-the-art at the time. REPA establishes a principled paradigm for improving representation learning in diffusion models through explicit cross-architecture semantic alignment.

Achieving state-of-the-art generation quality with fewer training steps.Enhancing training efficiency using external visual representations.Improving representation quality in diffusion models for generation.

Latest Papers

What's happening recently
View more

This work challenges the common assumption that different target representations are interchangeable in image generation, systematically investigating how representation choice affects generative difficulty. Within a unified masked autoregressive Rectified Flow framework, the authors evaluate raw pixels, SD-VAE latents, DINOv2 features, and MAE embeddings on ImageNet, analyzing their generation behavior and optimization characteristics. The study reveals that no single property—such as compressibility, reconstruction fidelity, dimensionality, or semantic clustering—sufficiently predicts generative performance. Instead, different representations redistribute difficulty across contextual modeling, per-token denoising, and control over the inference distribution: DINOv2 converges fastest but relies on a wide local denoiser, pixel-based training is slow and requires specialized configurations, MAE yields faithful reconstructions yet poor generative quality, and all representations exhibit markedly distinct trade-offs between precision and recall as well as responsiveness to guidance.

generative difficultyimage synthesismasked image generation

Existing image representation methods often struggle to simultaneously support both recognition and generation tasks. This work proposes a hypernetwork architecture based on Implicit Neural Representations (INRs), which encodes images into compact model weights that enable efficient reconstruction. By integrating knowledge distillation with pixel-level and perceptual losses, the method establishes a unified visual representation framework. It is the first approach to achieve high-accuracy recognition and high-quality image generation within a single shared embedding space, demonstrating state-of-the-art performance across diverse vision tasks while maintaining a highly compressed embedding dimensionality.

generationimage representation learningimplicit neural representation

Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing

Dec 19, 2025
SZ
Shilong Zhang
🏛️ The University of Hong Kong | Adobe Research | University of Chinese Academy of Sciences

Existing representation-based encoder generative paradigms face two key challenges: (1) discriminative feature spaces lack compact regularization, causing diffusion sampling to deviate from the data manifold and yield structural distortions; and (2) encoders exhibit weak pixel-level reconstruction capability, limiting geometric and textural fidelity. To address these, we propose a semantic-pixel joint reconstruction objective, achieving—within a compact 16×16, 96-dimensional latent space—the first unified high-semantic and high-fidelity pixel reconstruction. Our method integrates dual reconstruction losses, a compact latent-space design, a unified representation-encoder-based T2I and editing diffusion architecture, and a VAE feature-space adaptation mechanism. Experiments demonstrate significant improvements in reconstruction quality, text-to-image generation, and image editing—achieving state-of-the-art performance—along with accelerated convergence. This validates the feasibility of efficiently transferring understanding-oriented encoders into robust, generative latent spaces.

Adapting discriminative encoders for generative tasks with semantic-pixel regularizationEnabling compact, semantically rich latents for text-to-image generation and editingOvercoming off-manifold latents and weak reconstruction in representation encoders

Existing image tokenizers struggle to balance compactness with generation-friendliness. This work proposes a theory-driven regularization approach that, for the first time, incorporates frequency-aware dynamics from state space models into the image tokenization process. By designing a frequency-domain-aware regularization term, the method guides the tokenizer to learn latent representations that jointly capture spatial structure and spectral characteristics. The resulting tokenization preserves high reconstruction fidelity while significantly enhancing the generation quality of diffusion models, yielding a more efficient and generation-friendly image representation.

compact representationgeneration-friendlyimage tokenization

This work addresses the instability in representation alignment during diffusion model training, which arises from the mismatch between noisy inputs and clean image features, leading models to over-rely on complete token sets. To mitigate this alignment discrepancy, the authors propose MaskAlign, the first approach that operates from the perspective of token subsets. MaskAlign dynamically aligns representations using randomly masked token subsets and introduces a lightweight pre-mask token mixing module to encourage cross-token information sharing prior to masking. Integrated with a self-supervised visual encoder and diffusion Transformer training, MaskAlign significantly enhances generation quality and alignment robustness while maintaining computational efficiency, thereby improving the model’s generalization under perturbations of varying token subsets.

clean-image featuresdiffusion modelsnoisy inputs

Hot Scholars

LS

Linlin Shen

Shenzhen University
Deep LearningComputer VisionFacial Analysis/RecognitionMedical Image Analysis
SL

Simon Leglaive

CentraleSupélec - IETR (UMR CNRS 6164)
Speech and audio processingStatistical signal processingMachine and deep learning
BN

Bhalaji Nagarajan

Life Sciences Department, Barcelona Supercomputing Center
Deep LearningMachine LearningComputer Vision
IG

Imanol G. Estepa

Universitat de Barcelona
Self-supervised learningGenerative AI
SK

Suha Kwak

POSTECH
Computer VisionMachine Learning