Can 4D Foundation Models Remember?

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入PersistBench数据集和评估指标来解决4D基础模型记忆能力评估的问题,该方法能够以物体为中心的方式评价模型对物体永久性、运动连续性和外观保持的能力。
📝 Abstract
Perceiving and remembering the visual world is fundamental to navigating and interacting with our environment. Current 4D foundation models, such as camera-controllable video models or 4D reconstruction models, can perceive and reconstruct dynamic environments, but how well they remember what they have perceived remains an open question. Existing benchmarks largely rely on pixel-level metrics and lack ground truth for objects once they leave the field of view, making them unable to evaluate visual memory in an object-centric manner against references. To fill this gap, we introduce PersistBench, a dataset and metric suite that leverages 360° videos as omniscient ground truth and proposes three evaluation aspects: object permanence, motion continuity, and appearance preservation. Evaluating various models across diverse categories reveals that current models can only maintain short-term consistency that degrades significantly once objects leave the field of view. Our findings highlight the gap between current model capabilities and robust visual memory ("seeing is not remembering"), providing guidance for future development of 4D foundation models. Dataset and code are available on the project page: https://guangzhaohe.com/persistbench.
Problem

Research questions and friction points this paper is trying to address.

4D foundation models
visual memory
object permanence
motion continuity
appearance preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

PersistBench
visual memory
object permanence
motion continuity
appearance preservation