🤖 AI Summary
This study addresses the inability of current image-to-video models to ground behavior in mental state reasoning. To this end, we construct an evaluation benchmark employing zero-action prompts and a counterfactual experimental design. By integrating an automated evaluation pipeline with 744 counterfactual prompts and latent variable inference techniques, our approach effectively isolates psychological causal effects and identifies an “omniscience bias” failure mode. Evaluations across eleven mainstream models demonstrate that existing methods struggle to align with characters’ subjective beliefs, revealing a pronounced disconnect between visual fidelity and cognitive reasoning. This work bridges a critical gap in cognitive modeling for video generation and underscores the necessity of incorporating explicit theory-of-mind mechanisms into generative frameworks.
📝 Abstract
Current image-to-video models achieve visual realism and physical plausibility, but reasoning about mental states remains unexplored. Actions are driven by belief, desire, and perception, requiring inference beyond explicit instructions. We introduce MindWorldBench to evaluate mental-state-conditioned video generation. We formalize this as mental-state-to-behavior reasoning, where models generate actions from a world state and latent variables without explicit action prompts. MindWorldBench utilizes Zero-Action Prompting and a counterfactual design with 744 prompts to isolate the causal effects of mental states. An automated pipeline evaluates video quality, commonsense plausibility, and mental-state consistency. Evaluations of 11 models show that despite visual fidelity and physical reasoning, models fail to align behaviors with latent mental states. We identify a failure mode, termed Omniscient Bias, where models default to the objective world state rather than human's subjective belief. These results demonstrate a disconnect between visual generation and cognitive reasoning, suggesting a need for explicit mental-state modeling in video generation systems. Project website: https://richard2049-lee.github.io/MindWorldBench/