CARE: Condition-Aware Representation Regularization for Diffusion Models

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing regularization methods for diffusion models overlook the influence of conditioning signals on feature distributions, thereby limiting generation quality and efficiency. To address this limitation, this work proposes Condition-Aware Representation Regularization (CARE), a framework that dynamically modulates feature distributions via conditional similarity to encourage tight clustering of features under similar conditions. CARE operates without explicit alignment losses or external supervision, functioning as a plug-and-play module seamlessly compatible with existing regularization techniques. Experiments demonstrate that CARE reduces the FID by 19.08% and accelerates training by 3.5× on ImageNet. Furthermore, in text-to-image generation tasks, it achieves a 16.61% FID reduction alongside substantially improved semantic alignment.
📝 Abstract
Recent advances in diffusion models highlight the importance of representation regularization for improving sample quality and training efficiency. However, commonly used regularization methods often overlook the built-in conditions (such as labels or texts) which directly determine the generation target. In this work, we demonstrate how conditioning signals affect the feature distribution and introduce the CARE (Condition-Aware REpresentation regularization). CARE is a lightweight plug-and-play regularization framework that dynamically modulates feature distribution based on condition similarity. CARE leverages built-in conditioning signals to judiciously guide the representation space, promoting tighter feature clusters for similar conditions without relying on explicit alignment losses or external supervision. Empirically, CARE consistently improves both visual fidelity and convergence stability across both class-to-image and text-to-image tasks. On ImageNet, CARE achieves a 19.08\% reduction in FID in 400k training steps, leading to a 3.5$\times$ speed-up. When applied to text-to-image generation, CARE lowers FID by 16.61\% in 200k iterations and improves semantic alignment between generated samples and text prompts. Moreover, CARE can be seamlessly integrated with existing regularization methods, yielding additional performance gains.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Models
Representation Regularization
Condition-Aware
Sample Quality
Training Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Condition-Aware Representation Regularization
Diffusion Models
Plug-and-Play Framework
Feature Distribution Modulation
Text-to-Image Generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
F
Fengjia Guo
Department of Computer Science and Technology, Tsinghua University, Beijing, China
Zhuoyi Yang
Zhuoyi Yang
Tsinghua University
Deep Learning
Jie Tang
Jie Tang
UW Madison
Computed Tomography