What Visual Generators Need from Teachers: Rethinking Representation Alignment
This study addresses the reliance on empirical trial-and-error for teacher layer selection in knowledge distillation for diffusion models by proposing the RARE method. By revealing that student networks populate features in a bottom-up, layer-by-layer manner, this work defines a "recoverability gap" to replace conventional similarity metrics. This formulation enables adaptive alignment target selection and dynamic training loss weighting without trial-and-error. Evaluated on ImageNet, RARE reduces the FID to 4.46 with guidance, outperforming baselines such as REPA while decreasing GPU training time by 14%. Furthermore, the proposed approach demonstrates robust generalization across datasets of varying scales.