🤖 AI Summary
This study addresses the limitation of existing benchmarks in evaluating vision-language models (VLMs) on dynamic future state prediction and cross-view reasoning by proposing a novel visual-option-based paradigm for dynamic state reasoning. We construct a multi-source video evaluation benchmark comprising 2,503 questions and design four interconnected subtasks to systematically assess next-state prediction and bidirectional correspondence under first- and third-person perspectives. Extensive evaluations across proprietary, open-source, and spatial-reasoning VLMs reveal that the best-performing model achieves only 43.81% accuracy, substantially lagging behind human performance at 98.55%. These findings underscore significant limitations in current VLMs regarding spatiotemporal and cross-view compositional reasoning.
📝 Abstract
Vision-language models (VLMs) are increasingly evaluated for egocentric and cross-view video reasoning, yet existing benchmarks largely focus on semantic event understanding, temporal relations, or correspondence between already observed views, leaving their ability to reason directly about future visual states underexplored. We introduce EgoExo-Next, a visual-option benchmark for dynamic visual-state reasoning, where models must identify how an observed action trajectory subsequently appears rather than predict only an action label or textual description. EgoExo-Next contains 2,503 human-curated four-choice questions from six public egocentric and ego--exo video sources and comprises four interconnected subtasks that evaluate egocentric next-state prediction, bidirectional ego--exo state correspondence, exocentric next-state prediction, and their composition in Ego-to-Exo Next-State. Extensive evaluation of proprietary, open-source, and spatial reasoning VLMs reveals a substantial human--model gap, with the best model achieving 43.81\% average accuracy compared with 98.55\% for humans, and the largest degradation occurring on the composed Ego-to-Exo task. These results suggest that current VLMs remain substantially limited in dynamic visual-state reasoning, particularly when temporal progression and cross-view reasoning must be composed. The benchmark is publicly available at \url{https://huggingface.co/datasets/yutongli2024/EgoExo-Next}.