persistent object memory

Designs and implements mechanisms and data structures that maintain persistent representations of individual objects over long time intervals, enabling storage and retrieval of object identities, appearances, and internal states across occlusions, reappearances, and dynamic changes. Builds and analyzes indexing, update, and query methods for long-term object persistence to preserve continuity of object memory under temporal gaps and transformations.

persistentobjectmemory

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Video World Models with Long-term Spatial Memory

Jun 05, 2025
TW
Tong Wu
🏛️ Stanford University | Shanghai Jiao Tong University | Nanyang Technological University | The Chinese University of Hong Kong | Shanghai Artificial Intelligence Laboratory

Existing video world models suffer from limited temporal context windows, leading to environmental inconsistency and geometric forgetting upon scene revisitation. To address this, we propose the first video world model endowed with long-term spatial memory, leveraging a geometry-aligned 3D memory mechanism for cross-frame and cross-episode storage and retrieval of environmental states—thereby overcoming context-length bottlenecks. Our method integrates multi-view geometric modeling, memory-augmented Transformers, and a novel 3D memory encoding/retrieval module. We further introduce a dedicated long-horizon video-action paired dataset. Experiments demonstrate a 2.3× increase in effective context length, a 41% reduction in geometric error during scene re-visitation, and robust coherent generation over hundred-frame sequences. The model significantly outperforms baselines in both spatial-temporal consistency and visual fidelity.

Enhancing long-term scene consistency in video world modelsImproving quality and context length with 3D memory mechanismsReducing forgetting of previously generated environments

This paper addresses the challenge of simultaneously achieving stability and geometric fidelity in persistence diagram (PD) representation. We propose **Persistence Spheres**, a novel functional embedding method that constructs a **bicontinuous mapping** into a linear space—achieving, for the first time, **theoretically optimal preservation** of the 1-Wasserstein distance: the mapping is Lipschitz continuous with a Lipschitz continuous inverse. This guarantees topological stability while exactly recovering the geometric structure of PDs. The method admits an explicit, parallelizable formulation, enabling scalable topological data analysis. Across diverse domains—including functional data, time series, graphs, meshes, and point clouds—it consistently achieves or approaches state-of-the-art performance, significantly outperforming mainstream alternatives such as persistence images, persistence landscapes, and sliced Wasserstein kernels.

Creating bi-continuous functional representations of persistence diagramsEnsuring theoretical stability and geometric fidelity in representationsProviding computationally efficient persistence analysis for diverse data types

This work addresses the challenge of unreliable visual memory in interactive video world models during long-horizon generation, where positional encoding extrapolation failures and cache compression undermine visual persistence. The authors propose WorldTrace, a framework that—without requiring retraining—preserves cache addressability by assigning in-distribution virtual positions to compressed memory representations. WorldTrace supports two memory compression strategies: temporal coherence via WorldTrace-Field and event recall via WorldTrace-Landmark. This approach introduces the first training-free, addressable visual memory mechanism, effectively mitigating the phase conflict between RoPE extrapolation and compression. Evaluated on the newly introduced LoopBench benchmark, WorldTrace-Field improves temporal consistency by 15.5%, while WorldTrace-Landmark enhances event recall accuracy by 19.5%, significantly advancing long-range visually persistent generation.

KV cachememory compressionRoPE

This study addresses a critical limitation in continual learning research: the inability to distinguish whether conceptual information in model representations is genuinely lost or merely inaccessible. The work introduces, for the first time, a concept-level decomposition of forgetting by leveraging sparse autoencoders (SAEs) to construct a task-anchored latent space, framing forgetting along three dimensions—concept deletion, recoverability, and decodability. Through analysis of internal concept evolution in vision models, the authors demonstrate that most instances of “forgetting” stem not from irreversible information loss but from shifts in representational pathways that reduce decodability; under linear assumptions, the majority of concepts remain recoverable. This approach transcends conventional task-level evaluation paradigms and provides new insights into the fundamental mechanisms underlying forgetting in continual learning.

catastrophic forgettingconcept-level forgettingcontinual learning

This work addresses the limitations of existing interactive world models, which suffer from constrained spatial memory and insufficient 3D consistency due to the absence of explicit 3D environmental representations, thereby hindering long-term stable generation and downstream agent tasks. To overcome this, the authors propose PERSIST, a novel paradigm that integrates a persistent 3D state into world models for the first time. By combining implicit 3D scene modeling, differentiable rendering, and a dynamic state evolution mechanism, PERSIST jointly leverages spatial memory and user actions to generate videos. The method enables diverse yet geometrically consistent 3D environments from a single image and supports fine-grained, 3D-aware editing and control. Experiments demonstrate that PERSIST significantly outperforms current approaches in spatial memory, 3D consistency, and long-horizon stability, with both quantitative metrics and user studies confirming its superior generation quality.

3D consistencyinteractive generationpersistent state

Latest Papers

What's happening recently
View more

This work addresses the challenges of efficient retrieval, contextual selection, and evidential traceability in dynamic visual data streams by proposing a system-level abstraction that models visual memory as a multi-layered scene space aligned on a shared timeline. The design emphasizes source evidence traceability and decouples key decisions—such as model selection, segmentation, and embedding—into configurable components. Built upon the VideoDB format and a hierarchical indexing structure, the system treats real-time video streams as first-class data sources and offers typed search interfaces to support planned retrieval, stateful investigation, and evidence synthesis. Evaluated on four public datasets with over 9,800 queries, the generic component pipeline achieves Recall@1/@3/@10 of 73.09/83.39/91.20, significantly outperforming commercial video-native engines and demonstrating the critical impact of the proposed system architecture on retrieval performance.

continuous visual streamssource-grounded evidencetemporal granularity

Existing video editing methods struggle to maintain long-term semantic and structural consistency, primarily due to outdated contextual memory. This work proposes a decoupled multimodal context memory mechanism that constructs separate RGB and depth memory banks to model appearance semantics and geometric structure independently. By incorporating an edit-aware memory update and retrieval strategy, the method enables temporally and viewpoint-consistent video generation. Experimental results demonstrate that the proposed approach significantly outperforms current state-of-the-art techniques after editing, effectively preserving long-range semantic and structural coherence while exhibiting strong robustness.

consistent video generationcontext memorydisentangled representation

This work addresses the inefficiency of visual interface tokens in existing video large language models and their difficulty in explicitly modeling temporal associations among recurring objects. The authors propose SlotNarrative, a slot-based structured visual interface that parses videos into narrative sequences of persistent objects, representing each object and its temporal evolution through two compact token types: identity and state. SlotNarrative introduces a parameter-free memory mechanism that integrates multi-cue matching to cluster visual features into object slots, which are then fed into a frozen Video-LLM. Evaluated across multiple benchmarks, the method achieves superior accuracy and token efficiency using only 144 visual tokens, outperforming existing compact interfaces.

object persistencetemporal correspondencetoken efficiency

Hot Scholars

SS

Stefan Szeider

Professor, Head of Algorithms and Complexity Group, TU Wien, Vienna, Austria
AlgorithmsComplexitySatisfiabilityParameterized Complexity
NU

Naveed Ul Mustafa

Assistant Professor, CS, New Mexico State University
computer architecturememory security
MM

Miguel Matos

INESC-ID, Instituto Superior Técnico, Universidade de Lisboa
Distributed Systems