🤖 AI Summary
This study addresses the interference of generative history in evaluating multi-object memory within video world models by constructing an evaluation benchmark grounded in initial observations. Methodologically, it pioneers an input-anchoring mechanism and introduces explicit visibility reasoning. Combined with object-centric evaluators and dynamic-static trajectory tracking, this approach enables hierarchical quantitative measurement of existence, identity, and structural preservation capabilities. The findings reveal that retaining specific instances is significantly more challenging than generating plausible visual elements, and that model performance degrades systematically as the number of objects increases. These results provide critical empirical evidence for understanding the memory bottlenecks inherent in video world models.
📝 Abstract
Video world models must preserve the visual state of the world over time, but existing evaluation protocols often rely on generated histories, video reference, or selected revisit viewpoints that can confound the assessment of a model's true memory capability. To address this, we introduce OPIS, an input-grounded benchmark that strictly anchors the assessment to a fixed set of object instances from the initial observation for evaluating multi-object memory in video world models. The OPIS dataset comprises 500 cases across real-world, embodied-robotic, and game-world domains, providing dense object-level annotations for 12,672 rigid, articulated, and deformable instances. Our object-centric evaluator combines association and explicit visibility reasoning to hierarchically measure Object (O) Presence (P), Identity (I), and Structure (S), utilizing static or dynamic evaluation tracks based on object kinematics. Across eight image-to-video or camera-conditioned world models, our proposed OPIS scores range from 48.65 to 56.01. As the reference inventory grows from less than 20 to more than 40 objects, the Presence, Identity, and Structure scores show an overall decline, with the average Identity score falling from 40.22 to 23.11. The results demonstrate that preserving the particular object instances in the input is considerably harder than generating plausible visual elements.