STRIKE: Learning Visual State Transitions for Physical World Modeling
This study addresses the limitation of existing physical world modeling approaches, which predominantly focus on generating coherent motion while struggling to accurately capture interaction-induced scene state changes. To overcome this, we propose STRIKE, a framework that innovatively decouples visual state transition learning from dense video generation. Specifically, STRIKE employs an image-based transition model to predict subsequent scene configurations and integrates a pretrained vision-language planner to recursively generate future state sequences. A dynamics model then renders complete videos, with event-aligned supervision introduced to optimize training. By leveraging visual state transitions as intermediate representations, the proposed method effectively disentangles state prediction from video rendering. Extensive evaluations on benchmarks such as Physics-IQ Verified demonstrate that STRIKE significantly outperforms existing video backbone baselines in both physical consistency and manipulation fidelity.