Have I Scene This Before? Spatially Grounded Conversational Memory for Complex Queries in Egocentric Assistants

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency of first-person assistants in processing long-horizon spatial queries by proposing the SpaC-MEM framework. This method integrates 3D reconstruction with semantic segmentation to construct object-centric working memory, introducing a novel mechanism that anchors conversational information to persistent physical objects. Consequently, it enables multimodal history compression, spatial evidence preservation, and object-level fact updating, significantly reducing input token overhead. Experimental results demonstrate that SpaC-MEM achieves state-of-the-art accuracy and object recall on the Ego-SpaCR benchmark. Furthermore, it supports cross-scene and counterfactual spatial reasoning while maintaining substantially lower inference costs compared to native video baselines.
📝 Abstract
Egocentric assistants must connect what users say with what they see across long interaction histories. We formalize this challenge as Spatially grounded Conversational Reasoning (SpaCR): cross-scene, recall-oriented, and counterfactual spatial queries that combine user-stated facts with geometric evidence. Direct vision-language models incur high inference costs and context limits as histories grow, while keyframe selection and retrieval can omit objects or evidence needed for complete recall. We propose Spatially grounded Conversational Memory (SpaC-MEM), an object-centric working memory that uses 3D reconstruction and segmentation to ground conversational information in persistent physical objects. It compresses multimodal histories while preserving spatial evidence and allowing object-specific facts to be updated through dialogue. We also introduce Ego-SpaCR, a benchmark comprising 620 ScanNet video sessions augmented with 95 task-oriented conversations and 3,100 evaluation queries. SpaC-MEM achieves the highest overall answer accuracy among the evaluated methods and improves object recall while requiring substantially fewer reference input tokens than native-video baselines. Removing 3D spatial information substantially degrades performance, highlighting the importance of preserving spatial and conversational evidence together.
Problem

Research questions and friction points this paper is trying to address.

Egocentric assistants
Spatially grounded conversational reasoning
Cross-scene queries
Multimodal history compression
Object recall
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatially grounded Conversational Reasoning
Object-centric working memory
3D reconstruction
Egocentric assistants
Multimodal history compression
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jiazhou Liang
University of Toronto
L
Liam Gallagher
University of Toronto
K
Kiko Chen
University of Toronto
D
David Guo
University of Toronto
Armin Toroghi
Armin Toroghi
University of Toronto
LLMsReasoningNatural Language ProcessingKnowledge GraphsInformation Retrieval
Y
Yifan Simon Liu
University of Toronto
Scott Sanner
Scott Sanner
University of Toronto
Artificial IntelligenceMachine LearningInformation Retrieval