Keepsake: Selective Spatial Memory for Long-Horizon Video Generation

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high storage costs of retaining full history and the difficulty of dynamically evaluating frame importance in selective memory for long-horizon video generation. To overcome these challenges, we propose a training-free online memory controller that constructs a pose-appearance graph integrating camera pose proximity and visual similarity to reveal the relational nature of observational value. By continuously reconstructing the memory buffer under a fixed budget, the method dynamically retains irreplaceable frames while discarding redundant ones. Experiments demonstrate that our approach significantly reduces FVD and LPIPS metrics, achieving synergistic optimization of efficient memory management and generation quality. Notably, storing only 32 frames along a 180-second trajectory yields a 35.1% reduction in FVD, highlighting its effectiveness for long-duration synthesis.
📝 Abstract
Long-horizon camera-controlled video generation relies on persistent memory to maintain scene consistency. Existing systems follow two strategies to achieve this consistency. Full-history approaches retain all generated observations, causing unbounded storage and retrieval costs. Selective-construction approaches reduce redundancy, but make one-time retention decisions that are never revisited, even as an observation's value changes with the evolving memory bank. Both strategies leave a shared question unresolved: as the generated history evolves, which stored observations should still remain in memory? Our key insight is that the value of a stored observation is not fixed, but relational: it depends on the alternatives currently available in the memory bank. A view supported by many geometrically and visually similar substitutes can be relinquished with little loss of coverage, whereas an observation with few viable alternatives should remain regardless of age. We introduce Keepsake, an online, training-free controller for fixed-capacity spatial memory. At each update, Keepsake constructs a pose-appearance graph over retained and newly generated observations, combining camera-pose proximity with visual similarity. A retention priority jointly captures the number of strong substitutes and the similarity of the closest alternative, allowing Keepsake to continually reassess memory value, preserve observations with little alternative support, and evict highly replaceable ones under a fixed budget. The controller modifies only the persistent-memory update; the host generator, denoising schedule, and retrieval rule remain unchanged. Across MemCam and WorldMem, Keepsake improves FVD and LPIPS under a fixed memory budget. On 180-second MemCam trajectories, it retains only 32 of 5,397 frames while reducing FVD by 35.1%.
Problem

Research questions and friction points this paper is trying to address.

long-horizon video generation
spatial memory
scene consistency
memory management
camera-controlled generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Memory
Long-Horizon Video Generation
Pose-Appearance Graph
Training-Free Controller
Selective Retention