🤖 AI Summary
This study addresses the inability of existing video world model evaluations to directly assess the prediction of physical event consequences, particularly physical foresight at event boundaries. We propose an event-anchored evaluation framework grounded in real-world free-fall data, which for the first time decouples physical foresight into three independent dimensions: initiation, temporal progression, and motion dynamics. This framework is validated through multi-model comparisons and human baseline experiments. Our findings reveal that while mainstream models can trigger events, they frequently exhibit temporal lag or state freezing. Furthermore, we demonstrate that plausible temporal sequencing does not guarantee physically consistent motion. By precisely identifying these deficiencies, this work pinpoints the core bottlenecks limiting the physical foresight capabilities of current video world models.
📝 Abstract
Video world models are largely regarded as predictive models of the physical world and are therefore expected to anticipate the consequences of observed events. However, evaluation has mainly focused on reference similarity, physical-law consistency, or judgment plausibility, estimating anticipation only indirectly. We address this directly: when a release or impact has just occurred but its consequence is withheld, can a world model anticipate what should happen next? We introduce an event-anchored evaluation based on 62 controlled real-world free-fall recordings and 124 clips spanning three object types, with fine-grained release and impact annotations and ground-truth trajectories. The protocol separates consequence production, temporal placement, and physical realization. Across six contemporary video generation and world models, Runway and Veo produce release and subsequent impact events at rates above 93% but often initiate them substantially late, whereas Cosmos-Predict-2.5 and MAGI-1 frequently preserve the pre-event state and produce little or no measurable consequence. Among measurable falls, plausible timing does not necessarily imply physically consistent motion. We further conduct a 15-participant, 20-condition human study in which participants describe the expected consequence from a single event-anchored frame and draw its trajectory. Human predictions favor the recorded future in aggregate while revealing genuine ambiguity among plausible continuations. Overall, physical foresight emerges as a sequence of distinct challenges: initiating a consequence, anchoring it in time, and realizing its motion.