Towards Unified Evaluation of Prompt Enhancers for Video Generation

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing evaluations for video generation prompt enhancers, which rely on costly rendered videos and are prone to confounding with generation quality. To this end, this work proposes PEBench, the first unified benchmark designed to directly assess multimodal prompt enhancement capabilities. Methodologically, it constructs an evidence-oriented framework that integrates fact extraction with 24 criteria for fine-grained scoring, alongside an expert-validated case library, a task taxonomy, and an automated evaluation pipeline. The findings reveal that current prompt enhancement is evolving toward cinematic planning paradigms. Furthermore, PEBench scores demonstrate strong alignment with human judgments, establishing a reliable evaluation paradigm for future research in prompt enhancement.
📝 Abstract
Modern video generators can realize increasingly complex visual narratives, positioning the prompt enhancer (PE) as a critical bridge from concise user instructions and multimodal references to structured cinematic plans. However, existing PE evaluation relies on rendered videos, imposing substantial computational and human costs, slowing PE training and iteration, and conflating PE quality with downstream generator behavior. To address this gap, we introduce PEBench, the first unified benchmark for direct PE evaluation across text-to-video, image-to-video, and reference-to-video prompt enhancement. It comprises 1,100 expert-verified cases and 1,005 visual assets, spanning 35 fine-grained tasks with diverse temporal, cinematic, audiovisual, and multi-reference requirements. In addition, we develop PEBench evaluation, an evidence-grounded framework that combines modality-aware fact extraction with rubric-based assessment across 24 criteria. Our systematic evaluation of representative open- and closed-source PE methods reveals an emerging shift from fine-grained descriptive expansion toward intent-preserving cinematic planning, while the caption-reconstruction and forward-refinement methods show complementary strengths in cinematic coverage and semantic fidelity or internal coherence, respectively. Human validation shows that PEBench scores align closely with expert judgments of enhanced prompts and downstream videos from Wan3.0 and MiniMax-H3, indicating that prompt-level evaluation reliably reflects downstream utility.
Problem

Research questions and friction points this paper is trying to address.

Prompt Enhancer
Video Generation
Evaluation Benchmark
Prompt Enhancement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prompt Enhancer Evaluation
Video Generation Benchmark
Evidence-grounded Assessment
Modality-aware Fact Extraction
Cinematic Planning
Y
Yawen Shao
University of Science and Technology of China
Y
Yubo Zhu
Wan Team, Alibaba Group
Z
Ziyun Dai
Wan Team, Alibaba Group
Z
Zixun Fang
University of Science and Technology of China
K
Kai Zhu
Wan Team, Alibaba Group
Zeyinzi Jiang
Zeyinzi Jiang
Alibaba Group
Y
Yufeng Ai
Wan Team, Alibaba Group
Siyang Sun
Siyang Sun
Alibaba Group
deep learningmulti-modal large language model
H
Haolan Xue
Wan Team, Alibaba Group
Yu Shang
Yu Shang
Department of Electronic Engineering, Tsinghua University
Multimodal LearningLLM AgentRecommender System
Y
Yuxiang Bao
Wan Team, Alibaba Group
Z
Zoubin Bi
Wan Team, Alibaba Group
J
Jingming Luo
Wan Team, Alibaba Group
Jie Xiao
Jie Xiao
University of Science and Technology of China
low level visiongenerative modelmachine learning
Chaojie Mao
Chaojie Mao
Alibaba Group
Computer Vision
Zhehan Kan
Zhehan Kan
PhD student, Tsinghua University
CVMLLMsLLMs
H
Hongchen Luo
Northeastern University
Yu Liu
Yu Liu
Alibaba Group
self-supervised learninggenerative modeling
Sheng Zhong
Sheng Zhong
Nanjing University
computer networkssecurity and privacytheory of computing
Wei Tong
Wei Tong
School of Physics, University of Melbourne;
neural interfacemedical implantelectrophysiologybiomaterials
X
Xueyang Fu
University of Science and Technology of China
Yang Cao
Yang Cao
University of Science and Technology of China
computer visionimage processing
W
Wei Zhai
University of Science and Technology of China
Z
Zheng-Jun Zha
University of Science and Technology of China