🤖 AI Summary
This study addresses the challenge that vision-language models (VLMs) struggle to reason about the spatiotemporal states of objects moving out of view in first-person videos. To this end, we introduce Beyond3D, the first dynamic video question-answering benchmark designed for this purpose. Built upon the HD-EPIC dataset, Beyond3D leverages 3D object positions, camera poses, and scene geometry to construct object visibility trajectories, generating questions specifically targeting occlusion and displacement scenarios. This design isolates and evaluates VLM capabilities in tracking invisible objects, temporal localization, and spatial awareness. Experimental results demonstrate that the best-performing model achieves only 42.2% accuracy—significantly surpassing random baselines yet revealing critical deficiencies in recovering the last visible timestamps of objects. Ultimately, Beyond3D provides a vital evaluation tool and opens new research directions for out-of-view spatiotemporal reasoning in VLMs.
📝 Abstract
Real-world AI systems must reason about objects that are no longer visible: an AR assistant guiding a user back to an object used earlier, a household robot retrieving an item someone put away. This requires not just recalling where an object was last seen, but updating its state when it is moved and retaining that update once it leaves view. We refer to this as out-of-sight spatiotemporal reasoning. We introduce Beyond3D, the first VQA benchmark to isolate this ability in dynamic egocentric video: every query targets an object that has been relocated and has since left the field of view. We create our questions from HD-EPIC annotations, building a visibility track for each dynamic object from its 3D position, the camera pose, and the scene geometry to understand at each moment whether it is visible, occluded, or out of view. Beyond3D comprises 9,000 questions in eight types over 135 videos from nine participants, organized as one reasoning chain: visual grounding (is the target observable now), temporal grounding (when it was last visible and last placed), scene localization (which fixture anchors that location), and 3D spatial perception (where it lies relative to the current viewpoint or another object in the scene). We benchmark nine general-purpose and spatially specialized VLMs. The best model reaches 42.2% against 29.7% chance and text-only baselines reaching 31.9%, with the largest failures in recovering when an object was last visible, showing that tracking object movement out of sight remains far from solved for current VLMs.