🤖 AI Summary
This study addresses the limitations of existing evaluations for video generation prompt enhancers, which rely on costly rendered videos and are prone to confounding with generation quality. To this end, this work proposes PEBench, the first unified benchmark designed to directly assess multimodal prompt enhancement capabilities. Methodologically, it constructs an evidence-oriented framework that integrates fact extraction with 24 criteria for fine-grained scoring, alongside an expert-validated case library, a task taxonomy, and an automated evaluation pipeline. The findings reveal that current prompt enhancement is evolving toward cinematic planning paradigms. Furthermore, PEBench scores demonstrate strong alignment with human judgments, establishing a reliable evaluation paradigm for future research in prompt enhancement.
📝 Abstract
Modern video generators can realize increasingly complex visual narratives, positioning the prompt enhancer (PE) as a critical bridge from concise user instructions and multimodal references to structured cinematic plans. However, existing PE evaluation relies on rendered videos, imposing substantial computational and human costs, slowing PE training and iteration, and conflating PE quality with downstream generator behavior. To address this gap, we introduce PEBench, the first unified benchmark for direct PE evaluation across text-to-video, image-to-video, and reference-to-video prompt enhancement. It comprises 1,100 expert-verified cases and 1,005 visual assets, spanning 35 fine-grained tasks with diverse temporal, cinematic, audiovisual, and multi-reference requirements. In addition, we develop PEBench evaluation, an evidence-grounded framework that combines modality-aware fact extraction with rubric-based assessment across 24 criteria. Our systematic evaluation of representative open- and closed-source PE methods reveals an emerging shift from fine-grained descriptive expansion toward intent-preserving cinematic planning, while the caption-reconstruction and forward-refinement methods show complementary strengths in cinematic coverage and semantic fidelity or internal coherence, respectively. Human validation shows that PEBench scores align closely with expert judgments of enhanced prompts and downstream videos from Wan3.0 and MiniMax-H3, indicating that prompt-level evaluation reliably reflects downstream utility.