๐ค AI Summary
This study addresses the efficiency and accuracy bottlenecks in online distillation for few-step flow mapping generators, which stem from the coupling between student state acquisition and teacherโstudent distribution comparison. To this end, we propose a state-marginal-based decoupling framework. Specifically, we introduce a novel Rollout-Kernel separation mechanism that bridges local supervision with long-range mappings via flow-velocity consistency constraints, establishing an optimal strategy combining instantaneous velocity distribution supervision with tunable student consistency. Extensive experiments on ImageNet and text-to-image generation tasks demonstrate that our approach significantly outperforms GRPO baselines. Furthermore, it effectively enhances multi-expert integration capabilities, overall task performance, and convergence speed.
๐ Abstract
Few-step flow-map generators, including MeanFlow and consistency models, enable efficient sampling through long-range transport, yet their on-policy distillation remains underexplored. We introduce FlowMap-OPD, an on-policy distillation framework that separates student-state acquisition from teacher--student distribution comparison. A formulation based on state marginals establishes this separation, while flow--velocity consistency connects local supervision to the deployed long-range map. Within this framework, we develop flow-map, induced-velocity, and instantaneous-velocity distribution supervision, each paired with a separately specified native flow-map rollout. Cross-capacity ImageNet experiments across three teacher rewards identify instantaneous-velocity distribution supervision with independently tunable student consistency as the most effective choice. In text-to-image experiments, FlowMap-OPD demonstrates strong multi-specialist consolidation capabilities and surpasses multi-reward Flow-Map GRPO in task performance and convergence speed.