π€ AI Summary
Heterogeneous text-to-image diffusion models are difficult to integrate due to discrepancies in autoencoders and noise schedules, hindering on-demand utilization of multi-source capabilities within a unified framework. This work proposes a pixel-bridging strategy with an internal distillation framework that enables knowledge transfer from heterogeneous teacher models despite incompatible latent spaces. The approach employs pixel-level re-encoding, frozen DINOv2 for spatial alignment, shared attention LoRAs, and teacher-specific feedforward adapters. Furthermore, it introduces gradient compatibility diagnostics to guide adapter organization and gap-aware curriculum learning, enabling the first selective fusion of superior capabilities. Evaluated on a SD3.5-Medium (2.5B) student model, the method improves GenEval scores from 67.3 to 73.3 and DrawBench HPSv3 from 9.34 to 11.35, surpassing even larger teacher models.
π Abstract
Leading open text-to-image models often carry complementary strengths: one may lead on preference-aligned aesthetics while another follows compositional instructions more faithfully. However, differences in their autoencoders and noise schedules make it difficult to transfer these strengths across models. In this paper, we present Poly-OPD, a framework that can consolidate complementary strengths of heterogeneous teachers into a single compact flow-matching student. To bridge the incompatible latent spaces of different teachers, Poly-OPD performs on-policy distillation through a pixel bridge. Each student-generated image is re-encoded by a selected teacher's encoder and refined from a noise level matched by magnitude under the teacher's noise schedule. The resulting target is further matched to the student in frozen DINOv2 space, enabling supervision across incompatible latent spaces. To retain complementary capabilities without cross-teacher interference, Poly-OPD uses a gradient compatibility diagnostic to organize its adapters: attention LoRA modules are shared across teachers, whereas feed-forward adapters remain teacher-specific. During distillation, a gap-aware curriculum devotes more training to compositional categories where the student still falls short of the teacher. As each gap narrows, training shifts toward categories with larger remaining gaps. By distilling FLUX.1-dev and Z-Image into a 2.5B SD3.5-Medium student, Poly-OPD improves GenEval from 67.3 to 73.3, surpassing both larger teachers, and raises DrawBench HPSv3 from 9.34 to 11.35, consolidating both strengths within a switchable model.