How Does Geometry Enter Generated Motion?

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether the motion evolution in video generation models adheres to physical laws. To this end, it proposes a paired-intervention evaluation framework that conducts controlled experiments and physics benchmarking across nine image-to-video models by manipulating geometric variables such as trajectories and obstacles. The findings reveal that current models fundamentally perform geometry-conditioned motion synthesis, with their state evolution systematically deviating from physical laws. Although these models respond to geometric features that influence velocity, they lack the capacity for continuous physical state modeling. Furthermore, text prompts predominantly determine endpoints while duration governs pacing, and non-causal editing phenomena are observed. Collectively, this work offers novel insights into the physical limitations of video generation.
📝 Abstract
Under a fixed physical law, the visible geometry of a scene determines how motion must change. We ask how video generators realize this relationship. We fix the law and the initial state and change only the geometry drawn in the first frame, within matched families of tracks and deflectors, and compare each generated trajectory with the simulator prediction for that geometry. Paired interventions change one thing at a time: a local bump, the height of a barrier, the words of the prompt, the length of the clip. Across nine image-to-video models, geometry is preserved and shapes the motion: the speed of the ball follows the drawn undulation of a track. A physical state would carry this response forward, and here the generated motion parts from the law. The mean slope barely accelerates the ball, successive contacts fail to compose through a consistent state, an edit ahead of the ball alters its motion before it arrives, and the ball climbs over barriers higher than its release point. Two global conditions organize the global trajectory: text strongly controls the destination, while clip length strongly controls timing in the open-weight models tested. The pattern persists with photographed first frames. Current video generation thus behaves as geometry-conditioned motion synthesis whose evolution of state differs systematically from that of a fixed physical law.
Problem

Research questions and friction points this paper is trying to address.

video generation
physical law
geometry conditioning
motion synthesis
image-to-video models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Generation
Physical Reasoning
Causal Intervention
Geometry-Conditioned Motion
Mechanistic Interpretability
🔎 Similar Papers
No similar papers found.