π€ AI Summary
This study addresses the failure of vision-language models to correctly bind persons, actions, and spatiotemporal locations in videos by constructing a diagnostic benchmark comprising 6,701 video question-answering samples. Methodologically, it introduces a deterministic answer generation mechanism based on million-scale second-level annotations, designs five specialized tasks to evaluate spatiotemporal binding capabilities, and ensures data quality through multiple rounds of manual verification. Experiments reveal that mainstream models perform near random guessing on gaze detection and exhibit significant confusion regarding action agents. By systematically quantifying the spatiotemporal reasoning deficiencies of existing models, this work provides a high-quality evaluation framework and identifies promising directions for advancing video understanding research.
π Abstract
Video-capable vision-language models score above 80\% on popular benchmarks yet struggle with spatial-temporal binding: associating the right action with the right person at the right moment. We introduce ActionLens, a diagnostic benchmark of 6,701 multiple-choice video questions spanning five targeted diagnostics: transition detection, actor-specific identification, concurrent action binding, directed interaction reasoning, and gaze detection. Ground-truth answers are derived deterministically from 1.58 million per-second, per-person annotations. Fourteen rounds of human quality engineering raised answer clarity from 53% to above 90% human accuracy. Across 20 VLMs, the full-set leader scores 68.8%; on the human-reviewed subset, it scores 65.9% versus 91.0% for the pooled human reference. Gaze detection remains near chance against 89.6% human accuracy. On actor disambiguation, reference-interface controls show that relational descriptions recover 5.55--13.25 points over static coordinates, confirming a substantial numeric-parsing penalty; yet visual boxes still lead every model by 1.15--6.50 points, exposing a residual unboxed actor-resolution gap. A binding-trap analysis shows models systematically select the wrong actor's action. ActionLens provides diagnostic measurements of these distinct failure modes across model families and scales for direct comparison. We release all data, code, and evaluation scripts at https://anonymous.4open.science/r/lmms-eval-2276