VI-Bench: Benchmarking Prompt Inversion from AIGC Videos

📅 2026-09-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过建立VI-Bench解决了AIGC视频中提示词反转的问题,评估了18种模型在五个关键维度上的表现,揭示了现有模型的局限性。
📝 Abstract
Recent advances in video generation have made prompt-based control increasingly central to AIGC video generation. Prompts specify what a video should depict and how it should be represented, controlling factors such as visual style or camera behavior. Understanding this recoverability is important both for creative reuse and editing, and for assessing prompt leakage risks. However, existing video understanding benchmarks do not measure this capability: a caption may describe what is visible, but a replayable prompt must recover the generation-relevant controls needed to reproduce the video. To address this gap, we introduce VI-Bench, a benchmark built from 16.1 million real-user prompts and 900 human-verified AIGC videos. VI-Bench spans three progressively harder settings, namely single-shot semantic grounding, control over style and camera behavior, and multi-shot compositional inversion, and evaluates five generation-critical dimensions: subject, action, scene, style, and camera. We evaluate 18 representative VLMs, including 2 proprietary and 16 open-source models on VI-Bench, using an Inversion Score that measures prompt-level alignment with the original prompt and video-level fidelity of the regenerated video. The results reveal substantial limitations: even the strongest model achieves only 0.632 on Inversion Score, performance degrades sharply as samples require richer control and multi-shot reasoning, and models often produce plausible prompts whose regenerated videos deviate from the reference. These findings show that video prompt inversion is a distinct and under-evaluated capability requiring models to transform visual understanding into replay-stable generative control.
Problem

Research questions and friction points this paper is trying to address.

prompt inversion
video generation
benchmark
visual style
camera behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

VI-Bench
Prompt Inversion
AIGC Videos
Inversion Score
Generative Control
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Wulin Xie
Wulin Xie
Institute of Automation, Chinese Academy of Sciences
MLLMMulti-Modal
R
Rui Zhao
University of Virginia
Kecen Li
Kecen Li
Institute of Automation, Chinese Academy of Sciences
Data privacyMachine Learning
X
Xiujin Liu
University of Michigan, Ann Arbor
B
Bokang Zhang
The Chinese University of Hong Kong, Shenzhen
Z
Zheng Liu
University of Virginia
X
Xinwen Hou
Institute of Automation, Chinese Academy of Sciences
Chen Gong
Chen Gong
University of Virginia
PrivacyAI SecurityReinforcement LearningSoftware Engineering