WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of precisely planning long-duration, multi-shot cinematic sequences in text-to-video generation by proposing a director-level narrative model with 397 billion parameters. Methodologically, this work introduces video reverse engineering to automatically generate shot plans and proposes a semantic-consistent GRPO reinforcement learning algorithm. By integrating large language models with multimodal alignment techniques, the approach ensures instruction fidelity across shots and over extended temporal dimensions. Experimental results demonstrate that the proposed model significantly outperforms commercial counterparts in generating 5- to 15-second videos, while achieving substantial improvements in human preference scores for 30-second long-form videos, thereby establishing a new benchmark in the text-to-video generation domain.
📝 Abstract
Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfold across multi-shot sequences. In this paper, we present WanPE, a 397B-parameter prompt enhancement model trained on 1.05M real-world videos to master director-level cinematic planning. WanPE formulates shot-level cinematic plans via video-grounded reverse construction and employs Semantic-Consistency GRPO (SC-GRPO) to faithfully preserve user requirements across shots and over time. To benchmark this capability, we curate WanPEval, a human-annotated testbed covering durations from 5 to 30 seconds across varying intent granularities, supported by approximately 11K blind pairwise assessments. When powering Wan3.0's video generator, WanPE-397B boosts human preference over raw user prompts by 10.66-18.84 points at 5-15 seconds and by a dramatic 50.86 points in the 30-second arena. Ablation studies show that reverse construction demonstrates clear superiority over forward rewriting, while SC-GRPO robustly preserves semantic fidelity across model scales. Ultimately, WanPE leads all evaluated commercial offerings at 5-15 seconds and remains competitive with Seedance 2.5 at 30 seconds.
Problem

Research questions and friction points this paper is trying to address.

text-to-video generation
prompt enhancement
cinematic planning
multi-shot sequences
semantic consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prompt Enhancement
Reverse Construction
SC-GRPO
Text-to-Video Generation
Cinematic Planning
Y
Yubo Zhu
Nanjing University; Wan Team, Alibaba Group
Y
Yawen Shao
Wan Team, Alibaba Group; University of Science and Technology of China
Z
Ziyun Dai
Wan Team, Alibaba Group; Fudan University
Z
Zixun Fang
Wan Team, Alibaba Group; University of Science and Technology of China
K
Kai Zhu
Wan Team, Alibaba Group
Siyang Sun
Siyang Sun
Alibaba Group
deep learningmulti-modal large language model
H
Haolan Xue
Wan Team, Alibaba Group
Chuxin Wang
Chuxin Wang
University of Science and Technology of China
3D Computer Vision and 3D Object Detection
T
Tingyu Weng
Wan Team, Alibaba Group
J
Jingming Luo
Wan Team, Alibaba Group
C
Chen Shi
Wan Team, Alibaba Group
Lianghua Huang
Lianghua Huang
Tongyi Lab
generative modeling
Y
Yufeng Ai
Wan Team, Alibaba Group
Yuzheng Wang
Yuzheng Wang
Fudan University
Knowledge DistillationVision-Language ModelAIGC
W
Wenyuan Zhang
Wan Team, Alibaba Group
Yu Shang
Yu Shang
Department of Electronic Engineering, Tsinghua University
Multimodal LearningLLM AgentRecommender System
Y
Yuxiang Bao
Wan Team, Alibaba Group
Z
Zoubin Bi
Wan Team, Alibaba Group
Jie Xiao
Jie Xiao
University of Science and Technology of China
low level visiongenerative modelmachine learning
Jinbo Xing
Jinbo Xing
The Chinese University of Hong Kong
Computer Graphics and Vision
Jiaxing Zhao
Jiaxing Zhao
Jilin University
LLMsMulti-Agent
C
Chongyang Zhong
Wan Team, Alibaba Group
H
Hengjian Chen
Wan Team, Alibaba Group
C
Chenwei Xie
Wan Team, Alibaba Group
Akide Liu
Akide Liu
PhD Student @ Monash University
Efficient AIComputer Vision