Beyond Episodic Evaluation: Memory Architectural Bottlenecks in Sequential Embodied Question Answering

📅 2026-07-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of traditional Embodied Question Answering (EQA), which relies on episodic evaluation and struggles to support real-world robots in continuous tasks requiring accumulation and reuse of prior knowledge. Focusing on continuous, multi-turn EQA scenarios, the study systematically investigates the impact of memory architectures on embodied agent performance and proposes a structured, spatially anchored visual memory mechanism that maps persistent observations into a metric 3D geometric space. The approach reveals, for the first time, fundamental bottlenecks in existing memory designs and demonstrates that 3D geometry–based spatial anchoring is crucial for overcoming the trade-off between answer accuracy and navigation efficiency. Experiments show that the proposed architecture significantly improves accuracy while reducing navigation cost in simulation, and its effectiveness in continuous intelligent interaction is further validated on a real mobile robot.
📝 Abstract
Embodied question answering (EQA) is traditionally evaluated under an episodic formulation, where agents solve each task independently and reset internal state between episodes. However, real-world robots operate continuously and must accumulate, retain, and selectively reuse information acquired from prior interactions. Despite this practical requirement, the architectural mechanisms needed to support sequential memory in EQA remain underexplored. In this work, we investigate how different memory architectures behave when EQA agents are evaluated sequentially, with multiple questions answered in the same scene while memory is carried forward across queries. We find that simply preserving existing memory is often insufficient. Agents that retain only traversability information, such as 2D occupancy maps, remember where the robot has explored but not the visual-semantic evidence needed for later questions. Agents trained on short-horizon episodic data face a different challenge: when exposed to continuous, multi-query histories, their inherited context suffers from severe temporal mismatch, rather than forming a reusable scene representation. To overcome this architectural bottleneck, we highlight the necessity of structured, spatially grounded memory: architectures that map persistent visual observations onto metric 3D geometry preserve visual-semantic evidence in a coherent scene representation. Extensive experiments in simulated environments reveal that this form of memory breaks the accuracy-efficiency tradeoff in sequential settings, simultaneously achieving higher answer accuracy and lower navigation costs. We further validate these findings on a real-world mobile robot, demonstrating that spatially grounded visual memory is critical for enabling continuous, intelligent operation in physical environments.
Problem

Research questions and friction points this paper is trying to address.

Embodied Question Answering
Sequential Memory
Memory Architecture
Visual-Semantic Evidence
Scene Representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

structured memory
spatially grounded memory
sequential embodied question answering
3D scene representation
continuous operation
Zikui Cai
Zikui Cai
University of Maryland
Machine LearningTrustworthy AIComputer VisionRobotics
K
Kaushal Janga
University of Maryland, College Park
T
Tan Dat Dao
University of Maryland, College Park
S
Seungjae Lee
University of Maryland, College Park
Shivin Dass
Shivin Dass
PhD Student, UT Austin
Artificial IntelligenceMachine LearningRobot Learning
M
Mingyo Seo
The University of Texas at Austin
Kaiyu Yue
Kaiyu Yue
University of Maryland
Computer VisionMachine Learning
Mintong Kang
Mintong Kang
UIUC
Machine Learning
N
Nandhu Pillai
University of Maryland, College Park
M
Monte Hoover
University of Maryland, College Park
A
Aadi Palnitkar
University of Maryland, College Park
Ruchit Rawal
Ruchit Rawal
University of Maryland, College Park
InterpretabilityRobustnessMulti Modal Learning
Ruijie Zheng
Ruijie Zheng
University of Maryland, College Park, NVIDIA
Machine LearningReinforcement Learning
Bo Li
Bo Li
University of Illinois at Urbana–Champaign
Adversarial machine learningsecurityprivacybig datasocial network
Yuke Zhu
Yuke Zhu
The University of Texas at Austin, NVIDIA Research
Robot LearningComputer VisionMachine LearningRoboticsArtificial Intelligence
Roberto Martín-Martín
Roberto Martín-Martín
The University of Texas at Austin
RoboticsArtificial PerceptionMachine LearningInteractive PerceptionProbabilistic Reasoning
Tom Goldstein
Tom Goldstein
Volpi-Cupal Professor of Computer Science, University of Maryland
Numerical OptimizationMachine LearningDistributed ComputingComputer Vision
Furong Huang
Furong Huang
Associate Professor of Computer Science, University of Maryland
Trustworthy AI/MLReinforcement LearningGenerative AI