Learning to Retrieve: Internalizing Memory Retrieval for Video World Models

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing video world models rely on external memory systems, struggling to intrinsically learn when and what historical information to retrieve. This work proposes the L2R framework, which eliminates external memory banks by internalizing memory retrieval into the generation process. Specifically, this method achieves autonomous retrieval through a persistent internal state and a camera-conditioned gating mechanism, while introducing a 3D re-visibility signal to supervise the retrieval trigger, thereby making retrieval an inherent behavior of the model. Experimental results demonstrate that L2R significantly improves long-term scene consistency across multiple benchmarks. Furthermore, it operates efficiently during inference without requiring external memory or 3D conditioning.
📝 Abstract
Video world models aim to generate explorable, 3D-consistent scene videos conditioned on camera trajectories. Existing approaches often rely on external memory systems that explicitly retrieve previously observed content to mitigate scene drift during long-horizon generation. However, these auxiliary memory pathways operate outside the model's internal generative dynamics, preventing the model from intrinsically learning when and what historical information should be retrieved. We propose to internalize memory retrieval into the generation process, allowing retrieval to emerge as an intrinsic behavior of the video world model rather than relying on an external memory system. Based on this principle, we introduce \textbf{Learning-to-Retrieve (L2R)}, which repurposes the model's persistent internal state as a memory for historical context. A camera-conditioned retrieval gate selectively accesses relevant historical information from this state, determining \textit{what to retrieve}, while a retrieval trigger determines \textit{when to retrieve}. We further supervise the trigger with a 3D re-visibility signal, activating retrieval when previously observed content re-enters the current view while otherwise preserving the existing context. Together, these components enable the model to intrinsically acquire memory retrieval behavior and incorporate relevant historical observations into generation without a separate retrieval pathway. Across multiple base models and camera-revisit benchmarks, L2R improves long-term scene consistency while eliminating the need for an external memory bank or 3D conditions. https://jkhu29.github.io/l2r
Problem

Research questions and friction points this paper is trying to address.

Video World Models
Memory Retrieval
Scene Drift
Long-horizon Generation
3D Consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video World Models
Internalized Memory Retrieval
Learning-to-Retrieve
Retrieval Trigger
Long-term Scene Consistency
🔎 Similar Papers