transformer modulation for diffusion

Design and implement transformer-based modulation modules for diffusion generative models that adaptively scale and shift intermediate feature maps and inject conditioning signals. Build and evaluate DIT-style components that incorporate elapsed-time or auxiliary context into the denoising trajectory and calibrate generation to heteroscedastic noise.

transformermodulationfordiffusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.64
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Diffusion Models: A Comprehensive Survey of Methods and Applications

Sep 02, 2022
LY
Ling Yang
🏛️ Peking University | University of California, Los Angeles | Carnegie Mellon University | University of California, Merced

Despite their remarkable performance, diffusion models lack a systematic survey and a unified taxonomic framework. Method: This paper introduces the first comprehensive taxonomy encompassing methodological evolution and cross-domain applications, systematically reviewing over 300 seminal works published between 2015 and 2024. It focuses on three core research directions: efficient sampling, improved likelihood estimation, and modeling of structured data—covering key techniques including denoising score matching, stochastic differential equation (SDE) solvers, latent-space distillation, conditional guidance, and cross-modal joint modeling. Contribution/Results: We propose a novel integration paradigm that synergizes diffusion models with other generative paradigms. Furthermore, we publicly release a structured literature repository and a dynamically updated classification system, which has become a standard reference resource in the field.

Categorizing research into efficient sampling and likelihood estimationReviewing interdisciplinary applications across scientific fieldsSurveying diffusion models' methods and applications comprehensively

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the unclear mechanism of massive activations (MAs) in Diffusion Transformers (DiTs), which leads to insufficient generation detail and weak representational discriminability. The study reveals for the first time that MAs are spatially concentrated on image tokens and channel-wise focused on fixed dimensions, predominantly governed by the denoising timestep. Building on this insight, the authors propose EMA, a unified, training-free modulation framework that enhances generation quality through MA-driven detail guidance and boosts dense feature discriminability via MA-based modulation. Experiments demonstrate that EMA consistently improves both image generation fidelity and visual representation performance across diverse DiT architectures, offering strong local refinement capabilities while maintaining computational efficiency during inference.

Detail SynthesisDiffusion TransformersFeature Discrimination

Existing diffusion models often lack explicit awareness of image quality during the denoising process, leading to misaligned outputs, visual inconsistencies, and insufficient fidelity. To address this limitation, this work proposes a Quality Representation Module (QRM), which employs a lightweight Transformer to learn quality-aware representations conditioned on both textual prompts and timesteps. These representations modulate the adaptive LayerNorm layers within a Diffusion Transformer (DiT), thereby injecting quality-sensitive signals into the denoising dynamics. Notably, QRM introduces—for the first time—a lightweight quality-aware modulation mechanism into DiT architectures without altering the sampling strategy or backbone structure. Extensive experiments demonstrate that QRM consistently enhances image quality across multiple DiT baselines, and ablation studies confirm the effectiveness of its loss formulation and architectural design.

denoising processdiffusion transformersimage quality

To address high gradient variance, slow convergence, and reliance on normalization layers (e.g., AdaLN) in Diffusion Transformers (DiTs), this work proposes a magnitude-preserving network design and Rotation Modulation—a novel conditional modulation mechanism. The magnitude-preserving design replaces conventional normalization layers by constraining activation magnitudes, thereby enhancing training stability. Rotation Modulation parameterizes conditional transformations on the SO(2) group, substituting AdaLN’s scale-and-shift operations with lightweight, learnable 2D rotations. This is the first introduction of magnitude preservation into DiT architectures and the first rotation-based conditional modulation paradigm for diffusion models. Experiments demonstrate a 12.8% reduction in FID score; combining rotation modulation with scaling matches AdaLN’s performance while reducing parameter count by 5.4%. The implementation is open-sourced.

Improve performance and reduce parameters in diffusion modelsIntroduce rotation modulation as novel conditioning methodStabilize training in Diffusion Transformers via magnitude preservation

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Oct 09, 2024
SY
Sihyun Yu
🏛️ KAIST | Korea University | Scaled Foundations | New York University

Diffusion models suffer from inefficient representation learning and limited generation quality due to semantically impoverished latent spaces. To address this, we propose REPA (Representation Alignment), a novel regularization method that explicitly aligns denoising latent states—corrupted by noise—with clean-image representations extracted from high-quality external vision encoders (e.g., CLIP or DINO) within diffusion Transformers (DiT/SiT). This alignment is enforced via a projection-based loss, optimized end-to-end to enhance semantic consistency in the latent space. Experiments demonstrate that REPA accelerates SiT training by over 17.5×, enabling a SiT model to match the performance of a 7M-step SiT-XL within fewer than 400K steps. With classifier-free guidance (CFG), the method achieves an FID of 1.42—setting a new state-of-the-art at the time. REPA establishes a principled paradigm for improving representation learning in diffusion models through explicit cross-architecture semantic alignment.

Achieving state-of-the-art generation quality with fewer training steps.Enhancing training efficiency using external visual representations.Improving representation quality in diffusion models for generation.

Although Diffusion Transformers (DiTs) achieve impressive performance in image generation, their high computational and memory demands hinder deployment on edge devices. To address this challenge, this work proposes an efficient DiT framework featuring three key innovations: an adaptive global-local sparse attention mechanism to reduce computational complexity, an elastic training strategy within a unified hypernetwork enabling dynamic model scaling, and a four-step generative approach—KG-DMD—that integrates distribution matching with knowledge distillation. The resulting framework enables high-fidelity image generation in just four steps across diverse edge hardware platforms, significantly improving inference efficiency while maintaining visual quality and achieving a favorable balance between real-time performance and generation fidelity.

diffusion transformersedge devicesefficient deployment

Latest Papers

What's happening recently
View more

This work addresses the inefficiency in parameter utilization of Diffusion Transformers for generative tasks, which hampers their denoising performance. To overcome this limitation, the authors propose a lightweight calibration method that introduces approximately 100 learnable scaling parameters and formulates calibration as a black-box reward optimization problem, efficiently solved via an evolutionary algorithm. Requiring only minimal parameter fine-tuning, the approach significantly enhances generation quality and reduces inference steps across various text-to-image diffusion models while preserving high fidelity. This strategy achieves both parameter efficiency and computational efficiency in optimizing Diffusion Transformers, offering a practical and scalable solution for improving generative performance without extensive retraining or architectural modifications.

calibrationDiffusion Transformersgenerative quality

This work addresses the lack of theoretical understanding regarding why Transformers effectively learn optimal denoisers in diffusion models, particularly the mechanism by which they converge to the Bayes-optimal solution under non-convex loss. We establish, for the first time, a global convergence theory for Transformers trained on denoising diffusion probabilistic models (DDPMs) within a multi-label Gaussian mixture setting. By analyzing the population DDPM objective, modeling the multi-label Gaussian mixture distribution, and dissecting the self-attention architecture, we reveal how self-attention implements mean-field denoising and asymptotically approaches the minimum mean squared error (MMSE) estimator. Our analysis quantifies the required number of tokens per sample and training iterations to achieve a prescribed score-matching error. Numerical experiments corroborate the theoretical predictions and demonstrate alignment with MMSE estimation.

convergence analysisdenoisingdiffusion models

This work challenges the conventional assumption that diffusion models inherently require explicit timestep embeddings, investigating their necessity in the denoising process. Through theoretical analysis and empirical validation, the study demonstrates for the first time that under certain conditions, both U-Net and Diffusion Transformer architectures can converge to a global optimum without explicit timestep conditioning, implicitly inferring the noise scale. Ablation studies and generative evaluations on CelebA and CIFAR-10 show that such timestep-agnostic models achieve competitive or superior performance compared to standard timestep-conditioned counterparts in terms of FID, precision, and recall, while preserving high structural fidelity.

denoising processdiffusion modelsnoise scales

Existing methods struggle to accurately recover the initial noise latent variable from images generated by DDIM, achieving reasonable reconstruction quality but insufficient latent prediction accuracy. This work proposes a hybrid inversion approach that first employs gradient descent for direct inversion and subsequently refines the estimate through fixed-point iteration to more precisely recover the initial latent variable. The study introduces, innovatively, a “self-interpolation test” as a novel evaluation metric to comprehensively assess latent prediction fidelity. Experimental results demonstrate that the proposed method significantly improves both latent prediction accuracy and image reconstruction quality across three benchmark datasets, consistently outperforming existing approaches in self-interpolation test performance.

DDIM inversiondiffusion modelsimage generation inversion

Hot Scholars

HH

Haibin Huang

Principal Research Scientist at TeleAI
Computer GraphicsComputer VisionGeometric Modeling3D Deep Learning
PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
XY

Xiaoyan Yang

Advanced Digital Sciences Center
databasedeep learningtext mining
SY

Shengming Yin

University of Science and Technology of China
computer vision