FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of fine-grained evaluation frameworks grounded in professional cinematic language for assessing the artistic quality of generated videos. Collaborating with the Beijing Film Academy and industry practitioners, the authors construct the first text-to-video (T2V) and reference-to-video (R2V) benchmark that integrates filmic principles: they reverse-engineer authentic multi-shot prompts from award-winning films and propose a fine-grained evaluation framework comprising 3 axes, 12 components, and 38 sub-metrics. To enable automated scoring, they also release FilmOps, an open-source library of cinematic operation primitives. Experiments across 9 T2V and 7 R2V models demonstrate strong alignment between automatic scores and human judgments (Spearman’s ρ = 0.95/0.96), revealing significant deficiencies in current models regarding dynamic aesthetics and cross-shot consistency.
📝 Abstract
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they assess basic video plausibility rather than film-grade craft. We introduce FilmBench, a text-to-video (T2V) and reference-to-video (R2V) benchmark grounded in the professional Cinematic Language of the film- academy tradition and co-developed with directors and faculty from the Beijing Film Academy and the Hujing Digital Media & Entertainment Group film studio. It rests on three choices. First, prompts are reverse-engineered from clips of award-winning films spanning 20 cinematic genres and chosen by professional directors, so every prompt is anchored to a verified live-action reference; the prompts follow real shot lists, and most script multiple shots (1,056 of the 1,169 prompts are multi-shot), unlike prior single-clip benchmarks. Second, evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components and 35 (T2V) +3 (R2V-only) sub-metrics. Third, we develop an in-house expert-grade automatic evaluation agent and open-source its core suite of Cinematic Language operators (FilmOps). Benchmarking leading video generation models (9 for T2V, 7 for R2V), the evaluator reproduces the human model ranking at model-level Spearman \r{ho} = 0.95 (T2V) and 0.96 (R2V). Scores fall well below prior web-style benchmarks, with two consistent gaps in dynamic aesthetics and a marked single- to multi-shot performance drop that widens for weaker models.
Problem

Research questions and friction points this paper is trying to address.

video generation
cinematic evaluation
film-grade benchmark
Cinematic Language
professional filmmaking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cinematic Language
FilmBench
multi-shot prompting
expert-grade evaluation
FilmOps
🔎 Similar Papers
Shengyi Wang
Shengyi Wang
Princeton University
Programming LanguagesFormal Verification
N
Niantong Li
Alibaba Group
G
Guangzheng Hu
Alibaba Group
H
Hong Qi
Beijing Film Academy
Fei Ding
Fei Ding
Unknown affiliation
W
Weixu Qiao
Alibaba Group
Jinlin Wang
Jinlin Wang
DeepWisdom
Computer Vision、Multi-Agent System、Large Language Model、Large Vision-Language Model
X
Xiaotong Lv
Alibaba Group
Peng Han
Peng Han
Professor, Department of Computer Science, UESTC
drug discoveryspatial temporaldata mining
Z
Zimeng Li
Alibaba Group
F
Fanshu Ding
Alibaba Group
Y
Yushu Wang
Alibaba Group
H
Han Wu
Alibaba Group
Jingjing Chen
Jingjing Chen
Fudan University
MultimediaComputer VisionMachine LearningPattern recognition
C
Chongxiao Wang
Alibaba Group
Y
Yanhao Wu
Alibaba Group
C
Chenglong Huang
Alibaba Group
X
Xiaoqian Zhu
Alibaba Group
Jie Tian
Jie Tian
New Jersey Institute of Technology
Wireless Sensor NetworkAd hoc Sensor NetworkCloud Computing
H
Hua Li
Alibaba Group
J
Jingjing Fan
Alibaba Group
M
Mingshuang Tang
Beijing Film Academy
Z
Zhong Li
Beijing Film Academy
H
Hengxia Qiang
Beijing Film Academy
W
Weibin Chen
Alibaba Group