🤖 AI Summary
This study addresses the issue of poor instruction following in text-to-image models when given brief prompts, as well as the inference latency introduced by external prompt enhancers. To overcome these limitations, this work proposes PE-OPSD, which for the first time treats prompt enhancement as privileged training information. Built upon a flow matching framework, the method employs an enhanced-prompt teacher to perform online policy self-distillation over student trajectories, thereby internalizing the enhancement capability directly into the model. During inference, it processes raw prompts without requiring any additional components, completely eliminating external dependencies and latency overhead. Experimental results demonstrate that the proposed approach achieves state-of-the-art prompt fidelity and visual quality across multiple benchmarks while fully preserving the inference efficiency of the base model.
📝 Abstract
Text-to-image users often provide concise and underspecified prompts, whereas generative models benefit from detailed textual conditions for reliable instruction following. Existing systems bridge this gap with Prompt Enhancers (PEs) that rewrite raw prompts at inference time, introducing additional latency and leaving prompt elaboration external to the generator. We instead view enhanced prompts as privileged training information and ask whether their benefits can be internalized. We propose Prompt-Enhanced On-Policy Self-Distillation (PE-OPSD) for text-to-image flow-matching models. During training, a raw-prompt student follows its own generation trajectory, while an enhanced-prompt teacher provides vector-field targets at the states visited by the student. This on-policy supervision distills the behavior induced by enhanced prompts into the raw-prompt student without requiring additional text--image pairs. At inference, both the PE and teacher are removed, and the student generates directly from raw prompts. Across multiple model families, PEs, and benchmarks, PE-OPSD achieves the strongest aggregate prompt fidelity among the evaluated baselines, yields positive aggregate visual appeal gains, and retains the base-model inference efficiency.