🤖 AI Summary
This paper investigates the fundamental differences in representation learning mechanisms between diffusion models and classification models, specifically addressing whether diffusion models learn more balanced and comprehensive data representations. Method: We establish the first theoretical framework to comparatively analyze their feature learning dynamics, proving that the denoising objective implicitly optimizes feature balance and completeness. Combining theoretical analysis with experiments on both synthetic and real-world datasets—under standard U-Net architectures—we validate the mechanism via feature visualization and interpretability evaluation. Results: Diffusion models exhibit significantly enhanced inter-class separability and intra-class consistency in intermediate-layer features, outperforming structurally identical classification models. Our core contribution is uncovering the intrinsic regularizing effect of the denoising objective on representation quality, offering a novel theoretical perspective and empirical evidence for understanding the superior representational capacity of diffusion models.
📝 Abstract
The predominant success of diffusion models in generative modeling has spurred significant interest in understanding their theoretical foundations. In this work, we propose a feature learning framework aimed at analyzing and comparing the training dynamics of diffusion models with those of traditional classification models. Our theoretical analysis demonstrates that diffusion models, due to the denoising objective, are encouraged to learn more balanced and comprehensive representations of the data. In contrast, neural networks with a similar architecture trained for classification tend to prioritize learning specific patterns in the data, often focusing on easy-to-learn components. To support these theoretical insights, we conduct several experiments on both synthetic and real-world datasets, which empirically validate our findings and highlight the distinct feature learning dynamics in diffusion models compared to classification.