What Visual Generators Need from Teachers: Rethinking Representation Alignment

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance on empirical trial-and-error for teacher layer selection in knowledge distillation for diffusion models by proposing the RARE method. By revealing that student networks populate features in a bottom-up, layer-by-layer manner, this work defines a "recoverability gap" to replace conventional similarity metrics. This formulation enables adaptive alignment target selection and dynamic training loss weighting without trial-and-error. Evaluated on ImageNet, RARE reduces the FID to 4.46 with guidance, outperforming baselines such as REPA while decreasing GPU training time by 14%. Furthermore, the proposed approach demonstrates robust generalization across datasets of varying scales.
📝 Abstract
Representation alignment speeds up diffusion transformer training by pulling an intermediate block of the model (student) toward features of a frozen pretrained encoder (teacher). Which teacher layer to align, and for how long, is still set by convention, and each alternative costs a training run. We find that alignment helps where the student cannot linearly recover the teacher's features, not where it already resembles them. Since a deep teacher layer is largely predictable from the one below, we isolate what each layer adds, its increment, and measure how much of it an unaligned student recovers. The student fills the teacher's hierarchy from the bottom up and stalls near the top, which we call hierarchy filling: even after 400K steps it recovers almost none of the deepest. The recoverability gap is the unrecovered share of an increment, read from one unaligned checkpoint. In short runs that each align one teacher layer at one block, the gap nearly reproduces their ranking by FID improvement, and CKA, a measure of feature similarity, largely reverses it. Representation Alignment and Recoverability Estimation (RARE) picks the teacher layer with the largest gap before training. During training, it tracks each token's remaining distance to that layer, the online counterpart of the gap, weights tokens by it, and phases out the loss once the average distance stops falling. With SiT-B/2 on ImageNet $256\times256$, RARE reaches an FID of 18.02 without guidance and 4.46 with it, ahead of seven alignment baselines including REPA, iREPA and HASTE. It also trains in 14% fewer GPU-hours than iREPA. Its FID stays below iREPA's across model scales, teachers, datasets and backbones.
Problem

Research questions and friction points this paper is trying to address.

Representation Alignment
Diffusion Transformer
Teacher Layer Selection
Recoverability Gap
Visual Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Representation Alignment
Recoverability Gap
Diffusion Transformer
Hierarchy Filling
RARE
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yongcong Wang
Central South University
H
Hingchin Chen
The Hong Kong University of Science and Technology
M
Mingyu Fan
Tsinghua University
Shuo Jiang
Shuo Jiang
City University of Hong Kong
T
Teer Zhang
SenseTime Research
Y
Yucong Sun
SenseTime Research, Shandong University
Z
Zijia Wang
Imperial College London, University of Oxford, Dell Technologies
Y
Yiming Lu
University of International Relations
Chengchao Shen
Chengchao Shen
Central South University
Computer VisionMachine Learning