🤖 AI Summary
This study addresses the problem of implicit physical inconsistencies across different action predictions in action-conditioned video world models. To this end, it proposes a shared-mechanism counterfactual generation framework that jointly models multi-action future scenarios. Specifically, the method introduces a physical mechanism interpreter to infer latent physical distributions and employs a shared world evidence aggregation mechanism to unify underlying physical laws. Furthermore, an evidence-constrained flow matching training strategy is adopted to enforce physical consistency while preserving action specificity. Under a multi-intervention evaluation protocol, the proposed approach significantly enhances cross-intervention physical consistency while maintaining competitive single-rollout prediction quality.
📝 Abstract
Action-conditioned video world models aim to predict scene evolution under different actions, a capability that is essential for reliable planning, decision-making, and interaction in dynamic environments. However, futures generated independently from the same initial scene may each appear plausible while implying incompatible physical properties, such as friction or mass. This inconsistency can lead to contradictory predictions across interventions, making it difficult for the model to maintain a coherent understanding of the underlying world and limiting its reliability for planning and decision-making. To address these issues, we propose OneWorld, a shared-mechanism counterfactual generation framework that jointly models multiple action-conditioned futures under a common latent physical mechanism. A physical mechanism interpreter first infers a distribution over latent mechanisms from each action-outcome branch. These distributions are then aggregated into shared-world evidence, which captures whether the branches admit a common physical explanation while accounting for uncertainty in less informative branches. This evidence constrains flow training and guides sampling, encouraging consistency in the underlying physical mechanism while preserving the distinct outcomes induced by different actions. We further introduce a multi-intervention evaluation protocol in controlled environments, following the interaction settings of ACWM-Phys, to assess whether generated futures can be jointly explained by the same physical parameters, alongside standard measures of single-rollout prediction quality. Experiments in these environments show that OneWorld improves cross-intervention physical consistency while maintaining competitive single-rollout prediction quality.