🤖 AI Summary
This study addresses the limitation of existing physical world modeling approaches, which predominantly focus on generating coherent motion while struggling to accurately capture interaction-induced scene state changes. To overcome this, we propose STRIKE, a framework that innovatively decouples visual state transition learning from dense video generation. Specifically, STRIKE employs an image-based transition model to predict subsequent scene configurations and integrates a pretrained vision-language planner to recursively generate future state sequences. A dynamics model then renders complete videos, with event-aligned supervision introduced to optimize training. By leveraging visual state transitions as intermediate representations, the proposed method effectively disentangles state prediction from video rendering. Extensive evaluations on benchmarks such as Physics-IQ Verified demonstrate that STRIKE significantly outperforms existing video backbone baselines in both physical consistency and manipulation fidelity.
📝 Abstract
Physical world modeling requires predicting how interactions change a scene, not merely generating coherent motion. We propose STRIKE, a framework that separates visual state transition learning from dense video generation. We construct event-aligned supervision by extracting observed states from training videos and pairing them with transition descriptions and temporal offsets. An image-based transition model learns to predict the next scene configuration from the current image, a local transition specification, and elapsed time. At inference, a pretrained vision-language planner predicts time transition specifications, and recursive application of the learned transition model produces a sequence of future visual states. A separately trained dynamic model then generates the complete rollout conditioned on these states and their temporal locations. Experiments on Physics-IQ Verified, PhyGenBench, Pisa-Experiments, and RoboTwin2.0 show improvements of STRIKE over the corresponding video-backbone baselines in benchmark measures of physical consistency and manipulation-video fidelity. These results support learned visual state transitions as an effective intermediate representation for physical world modeling.