OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

📅 2026-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入OmniVBench和Omni-R2V数据集,解决了全方位参考到视频生成中基准测试不足及训练资源稀缺的问题,提供了更广泛的任务覆盖与细粒度控制。
📝 Abstract
Reference-to-video (R2V) generation is evolving toward increasingly general and versatile reference control, giving rise to the emerging paradigm of omni R2V generation. However, existing benchmarks fall short of these emerging capabilities: their test cases cover limited reference types and compositions, and their evaluation protocols largely assess holistic reference consistency, overlooking whether reference factors are properly preserved, disentangled, and routed. Meanwhile, the high cost of constructing omni R2V training data makes suitable training resources scarce. To address these gaps, we introduce OmniVBench and the Omni-R2V Dataset for evaluating and training omni R2V models. OmniVBench expands R2V evaluation across broader reference types, fine-grained control tasks, and richer reference compositions, covering 7 task families and 18 fine-grained tasks spanning content, motion, style, structure, narrative, and multi-reference settings. We introduce factor-grounded evaluation with 12,172 case-specific checklist items, assessing whether intended reference factors are faithfully preserved, correctly disentangled and bound to their targets, and properly realized according to the instruction. We further introduce the Omni-R2V Dataset, bringing industrial-grade training resources for diverse R2V tasks to the broader research community. Drawing primarily on a large-scale corpus of professional video footage, it comprises 340K processed training samples spanning diverse reference types and multi-reference compositions. We develop task-specific pipelines for reference-target pair construction, offering a practical and scalable recipe for omni R2V data construction. Extensive evaluation of advanced open- and closed-source R2V models reveals clear performance gaps across task families and evaluation dimensions on OmniVBench, highlighting remaining limitations of current R2V models.
Problem

Research questions and friction points this paper is trying to address.

reference-to-video generation
benchmark
evaluation protocol
training data
omni R2V
Innovation

Methods, ideas, or system contributions that make the work stand out.

OmniVBench
Omni-R2V Dataset
factor-grounded evaluation
multi-reference compositions
task-specific pipelines
🔎 Similar Papers
Wenxue Li
Wenxue Li
Harbin Institute of Technology Weihai
P
Peiyan Guan
Online Video BU, Tencent
H
Haoyang Jiang
Online Video BU, Tencent
J
Junxian Cai
Online Video BU, Tencent
H
Hualuo Liu
Online Video BU, Tencent
Chunjie Zhang
Chunjie Zhang
Beijing Jiaotong University
multimediacomputer vision
C
Chong Guan
Online Video BU, Tencent
S
Songlian Li
Online Video BU, Tencent
T
Taiyi Wu
Online Video BU, Tencent
Y
Yongjian Yu
Online Video BU, Tencent
X
Xiaotong Zhao
Online Video BU, Tencent
A
Alan Zhao
Online Video BU, Tencent
Eric Liu
Eric Liu
University of Toronto
SecurityCompilersFuzzing
X
Xi Chen
Online Video BU, Tencent
Y
Yu Liu
Online Video BU, Tencent
Lei Zhu
Lei Zhu
Hong Kong University of Science and Technology
Computational photographyVisionImage and video processingImage restorationhealthcare