OmniEcho: Spatial Audio Understanding for Embodied Agents

📅 2026-09-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为了解决实体代理在空间音频理解上的挑战,研究引入了OmniEchoBench基准和OmniEcho模型,通过结合视觉与音频信息提高场景推理和导航能力。
📝 Abstract
Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents. In particular, it is still unclear how to effectively evaluate and model spatial audio understanding in embodied settings. To address this gap, we introduce \textbf{OmniEchoBench}, a unified benchmark for spatial audio-visual perception and audio-vision-language navigation. OmniEchoBench comprises six tasks over 197 real-world spatial audio-visual scenes, 2,972 question-answer pairs, and 900 navigation samples with first-order ambisonics (FOA) audio collected from 30 real-world environments. To enable scalable training supervision, we develop a controllable rendering pipeline for spatial audio. It preserves geometric consistency among sound sources, visual observations, and agent trajectories. Building on this, we propose \textbf{OmniEcho}, a spatially aware omni-modal model. It introduces an FOA spatial encoder alongside a pretrained semantic audio pathway. Extensive experiments show that OmniEcho achieves state-of-the-art performance on spatial audio-visual perception. For our sound-guided navigation, OmniEcho reaches a performance level close to that of traditional vision-language navigation. These results demonstrate that spatial audio can serve as a valuable signal for embodied scene reasoning and navigation, while also highlighting fine-grained spatial localization and distance estimation as important open challenges.
Problem

Research questions and friction points this paper is trying to address.

spatial audio understanding
embodied agents
audio-visual perception
Innovation

Methods, ideas, or system contributions that make the work stand out.

OmniEchoBench
spatial audio understanding
first-order ambisonics (FOA)
OmniEcho
audio-visual perception
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Ruixun Liu
Ruixun Liu
Undergraduates of Xi'an Jiaotong University
computer vision
Y
Yuxuan Wang
Alibaba Token Hub, Alibaba Group
J
Jiacheng Xie
School of Intelligence Science and Technology, Peking University
Y
Yuhuan You
School of Intelligence Science and Technology, Peking University
D
Donghua Cai
Tsinghua University
J
Junming Lin
Tsinghua University
X
Xiong-Hui Chen
Alibaba Token Hub, Alibaba Group
Zhifang Guo
Zhifang Guo
Institute of Computing Technology Chinese Academy of Sciences
MultiModal/Speech/Sound/NLP
Yunfei Chu
Yunfei Chu
Alibaba Group
machine learning
Qize Yang
Qize Yang
Tongyi Lab, Alibaba Group
Computer VisionDeep Learning
X
Xize Cheng
Alibaba Token Hub, Alibaba Group
Jin Xu
Jin Xu
Qwen Team, Alibaba Group
Multimodal InteractionLarge Language ModelSpeech SynthesisVideo/Audio Processing
Yiwu Zhong
Yiwu Zhong
CUHK / University of Wisconsin-Madison
Vision-Language LearningMulti-Modal ModelsEmbodied AI