Representation by Design in Generation: Cross-View Class-Token Alignment in Diffusion Transformers

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that semantic representations in diffusion models emerge merely as byproducts of generation, making it difficult to enhance representational capacity without compromising generative quality. To overcome this, we propose a cross-view class-token alignment framework for diffusion Transformers that elevates representation learning to a primary optimization objective. The method introduces dual-timestep independent noise observations with an EMA teacher target, integrated with flow matching, self-supervised patch alignment, and stop-gradient mechanisms. We provide the first demonstration that diffusion model representations can be directly optimized rather than solely serving generation. Empirically, our approach yields substantial improvements in ImageNet linear probing accuracy and increases VOC segmentation mIoU by 3.6%, while maintaining FID scores and enhancing text-to-image synthesis quality, thereby achieving synergistic gains in both generation and representation.
📝 Abstract
Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas the models' own representations are often treated as a by-product of synthesis. We ask whether diffusion models can instead be trained to learn substantially stronger semantic representations without sacrificing generation quality. SelfFlow takes a step in this direction by introducing self-supervised patch alignment into flow matching, but its main gains remain in faster convergence and improved generation. Inspired by DINO and iBOT, we extend this framework with cross-view class-token alignment to further strengthen semantic representations. Specifically, we form two independently noised, dual-timestep observations of each image and align each student class-token representation with the stop-gradient EMA-teacher target from the other observation. This objective is optimized jointly with the inherited flow-matching and local patch objectives. Notably, although the additional objective acts only on the class token, it strengthens both class-token and patch representations. Compared with a matched two-view baseline, ImageNet linear-probing accuracy improves by 9.4\% using the class token and 10.1\% using mean-pooled patch tokens, while frozen-backbone VOC2012 segmentation improves by 3.6 mIoU. These representation gains are achieved while maintaining comparable ImageNet generation FID. In text-to-image training, the same objective also improves generation FID, reducing it from 2.52 to 2.37 at matched checkpoints. Our results show that representation need not remain a by-product of generation or merely a tool for improving it: it can be directly optimized as a first-class capability of diffusion pretraining alongside generation.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Models
Representation Learning
Semantic Representations
Generation Quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-View Class-Token Alignment
Diffusion Transformers
Self-Supervised Representation Learning
Flow Matching
EMA Teacher