FlowMap-OPD: Rollout--Kernel Separation for On-Policy Distillation of Few-Step Flow-Map Generators

๐Ÿ“… 2026-09-29
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the efficiency and accuracy bottlenecks in online distillation for few-step flow mapping generators, which stem from the coupling between student state acquisition and teacherโ€“student distribution comparison. To this end, we propose a state-marginal-based decoupling framework. Specifically, we introduce a novel Rollout-Kernel separation mechanism that bridges local supervision with long-range mappings via flow-velocity consistency constraints, establishing an optimal strategy combining instantaneous velocity distribution supervision with tunable student consistency. Extensive experiments on ImageNet and text-to-image generation tasks demonstrate that our approach significantly outperforms GRPO baselines. Furthermore, it effectively enhances multi-expert integration capabilities, overall task performance, and convergence speed.
๐Ÿ“ Abstract
Few-step flow-map generators, including MeanFlow and consistency models, enable efficient sampling through long-range transport, yet their on-policy distillation remains underexplored. We introduce FlowMap-OPD, an on-policy distillation framework that separates student-state acquisition from teacher--student distribution comparison. A formulation based on state marginals establishes this separation, while flow--velocity consistency connects local supervision to the deployed long-range map. Within this framework, we develop flow-map, induced-velocity, and instantaneous-velocity distribution supervision, each paired with a separately specified native flow-map rollout. Cross-capacity ImageNet experiments across three teacher rewards identify instantaneous-velocity distribution supervision with independently tunable student consistency as the most effective choice. In text-to-image experiments, FlowMap-OPD demonstrates strong multi-specialist consolidation capabilities and surpasses multi-reward Flow-Map GRPO in task performance and convergence speed.
Problem

Research questions and friction points this paper is trying to address.

on-policy distillation
few-step flow-map generators
consistency models
efficient sampling
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Few-Step Flow-Map Generators
Rollout-Kernel Separation
Instantaneous-Velocity Supervision
Consistency Models
๐Ÿ”Ž Similar Papers
No similar papers found.