RMS-AQA: A Two-Stage Spatial Audio Question Answering Benchmark for Real-World Domestic Environments

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the deficiency of embodied assistants in spatial audio localization and spatiotemporal reasoning within real-world domestic environments. To this end, it proposes a two-stage spatial audio question-answering benchmark that integrates first-order Ambisonics real-world recordings with high-fidelity synthetic data derived from measured room impulse responses. Furthermore, a lightweight spatial plugin is designed to enhance the spatial awareness of frozen audio-language models. This work systematically evaluates the full-pipeline capabilities of these models, ranging from sound source localization to complex question answering. The findings reveal that concurrent sound sources, long-distance propagation, and simulation-to-real domain shifts constitute the primary bottlenecks. Ultimately, this research provides a critical benchmark and identifies new directions for advancing spatial audio understanding.
📝 Abstract
Embodied assistants in domestic environments must infer what happened, where and when it occurred, and how to respond. To address this, we introduce RMS-AQA, a spatial audio question answering (SAQA) benchmark for real-world domestic environments. The benchmark features a two-stage question-answering (QA) format to comprehensively assess the ability of audio-language models (ALMs) to first ground audible sound events and subsequently perform complex spatio-temporal reasoning based on that grounding. To maximize acoustic realism, our dataset combines authentic real-world first-order Ambisonics (FOA) recordings with high-fidelity synthetic data generated using measured room impulse responses (RIRs). Furthermore, we provide a lightweight spatial plug-in that injects FOA-format data into frozen audio-language backbones. Experimental results reveal that the primary challenges stem from concurrent sources, far distance, and sim-to-real domain gap between RIR-synthesized and authentic recordings.
Problem

Research questions and friction points this paper is trying to address.

spatial audio question answering
embodied assistants
audio-language models
spatio-temporal reasoning
sim-to-real domain gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Audio Question Answering
Two-Stage QA
First-Order Ambisonics
Room Impulse Responses
Lightweight Spatial Plug-in
🔎 Similar Papers
2024-02-02International Conference on Machine LearningCitations: 14