Can Image Models Imagine Time? ImageTime: A Novel Benchmark for Probing Visual World Modeling Through Spatiotemporal Consistency

📅 2026-06-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing image generation models struggle to capture the temporal evolution of the visual world, often failing to maintain cross-frame identity, spatial relationships, and causal ordering. To address this limitation, this work introduces ImageTime, a benchmark that evaluates models’ ability to generate temporally coherent images under sequential instructions using a four-keyframe protocol—comprising initial, action-onset, transition, and final states—thereby abstracting away video-level dynamics to focus on logical consistency between static frames. We propose a novel diagnostic evaluation framework centered on spatiotemporal consistency as a probing mechanism, featuring a hierarchical task structure and structured state predicates. Leveraging a VLM-as-judge paradigm enables interpretable scoring and failure attribution. Through multi-stage state definitions, temporal constraint modeling, and causal violation detection—augmented by GPT-5.5–driven automated assessment—our approach systematically uncovers the capability boundaries, failure modes, and concept drift phenomena of state-of-the-art models in temporal visual consistency.
📝 Abstract
Image generation models now produce high-quality static images, yet their ability to represent how a visual world changes over time remains poorly understood. Practical workflows such as storyboarding, step-by-step illustration, reference-guided editing, and video previsualization require models to preserve identities, objects, spatial relations, and causal order across multiple visual states. Existing evaluations largely measure single-image correctness, compositional alignment, or video quality, leaving open whether an image model can coherently imagine a temporally ordered process. We introduce ImageTime, a diagnostic benchmark that uses spatiotemporal consistency as a behavioral probe of visual world modeling in image generation. Given an action instruction, and optionally a reference image specifying the initial state, a model must generate one image containing four ordered key states: initial state, action onset, transition state, and final state. This four-keyframe protocol is more temporally demanding than single-image generation while avoiding the confounds of dense video dynamics. ImageTime organizes tasks with a progressive capability hierarchy and decomposes each scenario into stage-wise state predicates, cross-frame temporal constraints, and forbidden causal violations. GPT-5.5 scores all generated images under a structured VLM-as-judge protocol, producing interpretable capability scores, diagnostic subscores, and failure labels. Through multi-family benchmarking, ImageTime reveals where current image generation systems succeed, fail, and drift when asked to maintain coherent visual world states over time.
Problem

Research questions and friction points this paper is trying to address.

image generation
temporal consistency
spatiotemporal modeling
visual world modeling
coherent time imagination
Innovation

Methods, ideas, or system contributions that make the work stand out.

spatiotemporal consistency
visual world modeling
temporal reasoning
image generation benchmark
VLM-as-judge