OPIS: An Input-Grounded Benchmark for Multi-Object Memory in Video World Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the interference of generative history in evaluating multi-object memory within video world models by constructing an evaluation benchmark grounded in initial observations. Methodologically, it pioneers an input-anchoring mechanism and introduces explicit visibility reasoning. Combined with object-centric evaluators and dynamic-static trajectory tracking, this approach enables hierarchical quantitative measurement of existence, identity, and structural preservation capabilities. The findings reveal that retaining specific instances is significantly more challenging than generating plausible visual elements, and that model performance degrades systematically as the number of objects increases. These results provide critical empirical evidence for understanding the memory bottlenecks inherent in video world models.
📝 Abstract
Video world models must preserve the visual state of the world over time, but existing evaluation protocols often rely on generated histories, video reference, or selected revisit viewpoints that can confound the assessment of a model's true memory capability. To address this, we introduce OPIS, an input-grounded benchmark that strictly anchors the assessment to a fixed set of object instances from the initial observation for evaluating multi-object memory in video world models. The OPIS dataset comprises 500 cases across real-world, embodied-robotic, and game-world domains, providing dense object-level annotations for 12,672 rigid, articulated, and deformable instances. Our object-centric evaluator combines association and explicit visibility reasoning to hierarchically measure Object (O) Presence (P), Identity (I), and Structure (S), utilizing static or dynamic evaluation tracks based on object kinematics. Across eight image-to-video or camera-conditioned world models, our proposed OPIS scores range from 48.65 to 56.01. As the reference inventory grows from less than 20 to more than 40 objects, the Presence, Identity, and Structure scores show an overall decline, with the average Identity score falling from 40.22 to 23.11. The results demonstrate that preserving the particular object instances in the input is considerably harder than generating plausible visual elements.
Problem

Research questions and friction points this paper is trying to address.

Video World Models
Multi-Object Memory
Benchmark Evaluation
Object Identity Preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video World Models
Input-Grounded Benchmark
Multi-Object Memory
Object-Centric Evaluator
Hierarchical Evaluation
🔎 Similar Papers
H
Hao Wang
USTC
T
Tao Yu
CASIA
L
Liuzhou Zhang
HKUST
H
HeXin Wang
Infrec
H
Haopeng Jin
CASIA
Y
Yuxuan Zhou
Tsinghua University
X
Xinming Wang
CASIA
H
Hongzhu Yi
UCAS
Xinye Li
Xinye Li
Harbin Institute of Technology
Large Language ModelNatural Language ProcessingAgentKnowledge EditingInterpretability
Y
Yuanlei Wang
Sun Yat-sen University
Ping Nie
Ping Nie
Waterloo University
Natural Language ProcessingInformation RetrievalRecommendation SystemsTime Series Forecasting
Y
Yan Huang
CASIA
Y
Yuxuan Zhang
Jiangnan University
P
Pengfei Zhou
Infrec
Y
Yanyan Zou
SUTD
W
Wei Yang
USTC