DataVista: Diagnosing Multimodal LLMs on Data Video Understanding

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic evaluation benchmarks for data video understanding and the difficulty existing multimodal models face in accurately parsing dynamic charts and integrating narrative evidence. We introduce the first benchmark for data video understanding, comprising 961 real-world videos and 6,775 questions, alongside a three-tier progressive capability framework encompassing perception, temporal reasoning, and narrative understanding across ten fine-grained question-answering tasks. The evaluation integrates video frame sampling, caption augmentation, and automated question generation to comprehensively assess mainstream multimodal large language models. Experimental results reveal that the best-performing model achieves only 70% accuracy, substantially trailing human experts and exposing critical deficiencies in causal reasoning and narrative structure modeling.
📝 Abstract
Data video is a media form that integrates data visualization with video narrative, widely adopted in news reporting and business analysis. Compared with general video understanding, data video understanding places greater emphasis on accurately reading data from animated charts, integrating evidence across charts and time, and understanding how narrative organization and visual design communicate information. Yet existing benchmarks target either general videos or static charts, and data video understanding has not been systematically evaluated. We present DataVista, the first benchmark for data video understanding, containing 961 real-world data videos and 6,775 evaluation questions organized under a three-level progressive capability framework (data perception, temporal reasoning, narrative understanding) with 10 fine-grained question types across five topic domains. Systematic evaluation of 19 mainstream MLLMs shows that the best-performing model, Gemini-3.1-Pro, achieves 70.0% overall accuracy, still far below human expert performance, with models performing worst on Causal Reasoning and Narrative Structure. Increasing frame counts and adding subtitles mainly benefit data perception and temporal reasoning, with limited gains in narrative understanding. Further analysis of model responses identifies typical failure modes in chart reading, evidence judgment, and instruction understanding. The benchmark is available at https://github.com/HKUSTDial/DataVista.
Problem

Research questions and friction points this paper is trying to address.

Data Video Understanding
Multimodal LLMs
Benchmark Evaluation
Dynamic Charts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data Video Understanding
Multimodal LLMs Benchmark
Progressive Capability Framework
Temporal Reasoning
Narrative Understanding
💼 Related Jobs
No related jobs found.
Y
Yupeng Xie
The Hong Kong University of Science and Technology (Guangzhou)
Z
Zhenyang Wang
The Hong Kong University of Science and Technology (Guangzhou)
Jiayi Zhu
Jiayi Zhu
Ph.D student, state key laboratory of cognitive neuroscience and learning, Beijing Normal University
Cognitive neuroscienceNeuroimagingDeep learning
Yinghao Tang
Yinghao Tang
State Key Lab of CAD&CG, Zhejiang University
Large Language ModelMLSystem
Z
Zhouan Shen
The Hong Kong University of Science and Technology (Guangzhou)
Y
Yiyu Chen
The Hong Kong University of Science and Technology (Guangzhou)
Yuyu Luo
Yuyu Luo
Assistant Professor, HKUST(GZ) / HKUST
Data AgentsLLM AgentsDatabaseText-to-SQLData-centric AI