Informative Viewpoint Selection for Episodic-Memory Embodied Question Answering using Omnidirectional Images

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of severe geometric distortion, excessive background noise, and visual token redundancy in omnidirectional images for embodied question answering. Focusing on open-vocabulary episodic memory tasks, this work proposes a cube-projection-based viewpoint selection method. Specifically, it eliminates distortion through equirectangular-to-perspective projection, evaluates image-text relevance using a fine-tuned BLIP-2 model, and selects an optimal subset of viewpoints via a diversity-aware greedy algorithm to enhance input quality for vision-language models. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance on the HM3D dataset, reducing the number of observed frames by 65.5% while maintaining high answer accuracy.
📝 Abstract
Embodied Question Answering (EQA) requires agents to answer natural language questions about surrounding environments from visual observations. In this work, we focus on open-vocabulary episodic-memory EQA (EM-EQA), where an agent answers free-form questions using recorded observation histories. Omnidirectional images are promising for this task, as they provide wide field-of-view observations that can capture surrounding context without requiring explicit camera rotations. However, omnidirectional images introduce two challenges for EQA: (i) equirectangular projection causes severe geometric distortion that degrades vision-language model (VLM) recognition accuracy, and (ii) feeding equirectangular images directly into VLMs introduces excessive irrelevant background information, reducing answer accuracy and increasing the visual-token burden. To address these challenges, we propose a viewpoint selection method for EM-EQA using omnidirectional images. Our method converts equirectangular observations into perspective views via cubemap projection, estimates question-conditioned relevance with fine-tuned BLIP-2, and selects informative and diverse viewpoints through diversity-aware greedy selection. Experiments on the Habitat-Matterport 3D (HM3D) subset of OpenEQA show that our method achieves state-of-the-art model performance among the reported model results with equirectangular observations. Moreover, after removing rotation views, which reduces observation frames by 65.5%, our method largely maintains its answer accuracy.
Problem

Research questions and friction points this paper is trying to address.

Embodied Question Answering
Omnidirectional Images
Episodic Memory
Equirectangular Projection
Vision-Language Model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Episodic-Memory Embodied Question Answering
Omnidirectional Images
Viewpoint Selection
Cubemap Projection
BLIP-2
💼 Related Jobs
No related jobs found.
K
Kaname Kitamura
Department of Computer Science, Institute of Science Tokyo, Tokyo, Japan
Asako Kanezaki
Asako Kanezaki
Tokyo Institute of Technology
Computer VisionObject RecognitionShape Matching