🤖 AI Summary
This work addresses the challenge of robotic manipulation in visually constrained environments, where existing methods struggle to effectively leverage auditory cues for target localization and identification. To this end, we propose the first multimodal audio-visual imitation learning framework that jointly integrates acoustic spatial location, timbral characteristics, and visual information, supporting diverse policy architectures including ACT, Diffusion Policy, VQ-BeT, and π₀. Our approach uniquely incorporates both raw audio and acoustic spatial cues into imitation learning and introduces a new benchmark for acoustically grounded manipulation tasks. Experimental results demonstrate that the proposed method achieves state-of-the-art performance on tasks requiring joint reasoning about sound location and timbre, validating its effectiveness across both simulated and real-world robotic platforms.
📝 Abstract
Acoustic information provides rich cues about object location, material properties, and changes caused by contact or motion. This paper introduces a new set of acoustic-aware manipulation tasks for imitation learning, in which robots must use auditory cues to determine manipulation targets. These tasks require sound source localization and identification for active exploration in robotic manipulation. Also, we propose a multimodal imitation learning framework, Spatial-Spectral Audio Action (S2A2), that integrates visual features with acoustic spatial and acoustic signal information for the acoustic-aware manipulation tasks. We implemented S2A2 models that integrates policies such as ACT, Diffusion Policy, VQ-BeT, and $π_0$, into our framework. Simulation experiments showed that the proposed method is the most effective for tasks requiring both position and timbre. Furthermore, real-robot experiments confirm the applicability of the proposed tasks and framework to real-world manipulation.