🤖 AI Summary
This study addresses the low training efficiency of Diffusion Transformers (DiTs) and the limitations of existing representation alignment methods by proposing the REPI framework. Adopting an approach inversely complementary to REPA, REPI employs a "scaffolding-to-internalization" strategy that injects representations from pretrained vision encoders into DiTs, enabling their deep involvement in the denoising process for progressive internalization. Notably, REPI can be combined with REPA to yield synergistic gains. Experimental results demonstrate that this method surpasses the performance of a 7-million-step baseline using only 160,000 training steps, achieving over a 43.5× speedup. By significantly outperforming mainstream approaches, REPI establishes a new paradigm for efficient DiT training.
📝 Abstract
Recent representation alignment (REPA) methods accelerate diffusion transformer training by aligning projections of the transformer's hidden states with representations from pretrained visual encoders. In this work, we explore a reverse and complementary direction to REPA: rather than projecting diffusion representations into the encoder's space, we inject encoder representations into the diffusion transformer, allowing them to actively participate in the denoising process. To this end, we introduce \textit{REPresentation Injection} (REPI), a training framework based on a scaffold-to-internalization strategy, in which projected encoder representations initially serve as a temporary scaffold and are then progressively internalized by the diffusion transformer. REPI outperforms REPA across a wide range of backbones and is highly complementary to it: combining the two yields substantial gains over either alone. Notably, with only 160K training steps, REPI + REPA matches vanilla SiT trained for 7M steps, a speedup of over $43.5\times$. Code will be available at https://jeneveuxpas.github.io/REPI