π€ AI Summary
This work challenges the conventional assumption that diffusion models inherently require explicit timestep embeddings, investigating their necessity in the denoising process. Through theoretical analysis and empirical validation, the study demonstrates for the first time that under certain conditions, both U-Net and Diffusion Transformer architectures can converge to a global optimum without explicit timestep conditioning, implicitly inferring the noise scale. Ablation studies and generative evaluations on CelebA and CIFAR-10 show that such timestep-agnostic models achieve competitive or superior performance compared to standard timestep-conditioned counterparts in terms of FID, precision, and recall, while preserving high structural fidelity.
π Abstract
Diffusion models rely heavily on explicit timestep embeddings to modulate the denoising process across various noise scales. In this work, we challenge the necessity of these temporal signals by analyzing their impact on U-Net and Diffusion Transformer architectures. Beyond empirical evidence, we provide a theoretical framework demonstrating that, under certain conditions, the global minimizer of the diffusion training objective can be achieved without explicit timestep conditioning. Our findings reveal a surprising robustness when timestep embeddings are completely removed. Extensive ablation studies on the CelebA and CIFAR-10 datasets show that these time-agnostic models can maintain high structural fidelity and even surpass their conditioned counterparts in competitive metrics, including FID, precision, and recall. Our analysis suggests these architectures can implicitly infer noise scales from the corrupted input under specific assumptions, rendering explicit temporal conditioning redundant. This study challenges long-standing temporal conditioning paradigms and paves the way for more efficient and structurally focused generative architectures.