🤖 AI Summary
This study addresses the deficiency of embodied assistants in spatial audio localization and spatiotemporal reasoning within real-world domestic environments. To this end, it proposes a two-stage spatial audio question-answering benchmark that integrates first-order Ambisonics real-world recordings with high-fidelity synthetic data derived from measured room impulse responses. Furthermore, a lightweight spatial plugin is designed to enhance the spatial awareness of frozen audio-language models. This work systematically evaluates the full-pipeline capabilities of these models, ranging from sound source localization to complex question answering. The findings reveal that concurrent sound sources, long-distance propagation, and simulation-to-real domain shifts constitute the primary bottlenecks. Ultimately, this research provides a critical benchmark and identifies new directions for advancing spatial audio understanding.
📝 Abstract
Embodied assistants in domestic environments must infer what happened, where and when it occurred, and how to respond. To address this, we introduce RMS-AQA, a spatial audio question answering (SAQA) benchmark for real-world domestic environments. The benchmark features a two-stage question-answering (QA) format to comprehensively assess the ability of audio-language models (ALMs) to first ground audible sound events and subsequently perform complex spatio-temporal reasoning based on that grounding. To maximize acoustic realism, our dataset combines authentic real-world first-order Ambisonics (FOA) recordings with high-fidelity synthetic data generated using measured room impulse responses (RIRs). Furthermore, we provide a lightweight spatial plug-in that injects FOA-format data into frozen audio-language backbones. Experimental results reveal that the primary challenges stem from concurrent sources, far distance, and sim-to-real domain gap between RIR-synthesized and authentic recordings.