🤖 AI Summary
Current evaluation of Video Large Language Models (VideoLLMs) suffers from fragmented benchmarks, inconsistent evaluation protocols, data leakage, and generalization bias. Method: We systematically survey mainstream video understanding benchmarks and present the first comprehensive taxonomy of VideoLLM evaluation methodologies. Through benchmark analysis, protocol categorization, performance trend statistics, and limitation diagnosis, we identify critical bottlenecks across closed-set, open-set, and spatiotemporal understanding paradigms. We propose next-generation benchmark design principles centered on diversity, multimodal alignment, and interpretability, and develop a framework evolution model to characterize performance patterns of leading VideoLLMs across benchmarks. Contribution/Results: This work delivers the field’s first structured, principled evaluation guide—enabling standardized, rigorous, and scientifically grounded assessment of VideoLLMs—and advances the maturation of video understanding evaluation.
📝 Abstract
The rapid development of Large Language Models (LLMs) has catalyzed significant advancements in video understanding technologies. This survey provides a comprehensive analysis of benchmarks and evaluation methodologies specifically designed or used for Video Large Language Models (VideoLLMs). We examine the current landscape of video understanding benchmarks, discussing their characteristics, evaluation protocols, and limitations. The paper analyzes various evaluation methodologies, including closed-set, open-set, and specialized evaluations for temporal and spatiotemporal understanding tasks. We highlight the performance trends of state-of-the-art VideoLLMs across these benchmarks and identify key challenges in current evaluation frameworks. Additionally, we propose future research directions to enhance benchmark design, evaluation metrics, and protocols, including the need for more diverse, multimodal, and interpretability-focused benchmarks. This survey aims to equip researchers with a structured understanding of how to effectively evaluate VideoLLMs and identify promising avenues for advancing the field of video understanding with large language models.