VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决视频大语言模型细粒度理解评估难题,通过构建VidOmni-Bench基准,要求模型验证密集视频字幕中的每个事件是否得到视频支持。
📝 Abstract
While Video Large Language Models (Video-LLMs) have recently demonstrated strong performance, reliably evaluating their fine-grained video understanding remains challenging. Existing benchmarks often rely on question answering or ground-truth caption matching, where models may succeed through superficial cues and incomplete annotations. To this end, we introduce VidOmni-Bench, a benchmark that requires models to verify whether each event in dense video captions is supported by the video. VidOmni-Bench consists of 500 videos spanning five complexity types and diverse durations from 4 seconds to 90 minutes. After collecting videos along these axes, we use diverse Video-LLMs to generate dense captions and obtain human-verified sentence-level labels, where sentences containing incorrect events serve as hard negatives for evaluation. Our experiments on VidOmni-Bench reveal three key findings: (i) Video-LLMs frequently generate hallucinated descriptions in dense video captioning; (ii) they also struggle as verifiers, failing to reliably detect plausible but incorrect event descriptions; and (iii) model weaknesses vary across video complexity and duration, revealing diverse, model-specific bottlenecks in current Video-LLMs.
Problem

Research questions and friction points this paper is trying to address.

Video-LLMs
fine-grained video understanding
benchmark
event verification
dense captions
Innovation

Methods, ideas, or system contributions that make the work stand out.

VidOmni-Bench
fine-grained video understanding
spatio-temporal event verification
dense video captions
human-verified labels
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30