🤖 AI Summary
This work addresses the neglect of acoustic cues during object impact in existing 3D reconstruction methods and the limitations of traditional sound modeling, which relies on costly physical simulations or large datasets and struggles to generalize under few-shot conditions. To overcome these challenges, the authors propose AV-MSF, a novel object-level acoustic field representation that uniquely integrates 3D Gaussian splatting with dense visual features, incorporating geometry-aware priors and physically interpretable modal parameters. This approach enables high-fidelity impact sound reconstruction from limited examples. Evaluated on two real-world datasets, AV-MSF significantly outperforms both physics-based simulation and purely data-driven baselines, while also demonstrating strong performance in downstream tasks such as contact localization and sound editing.
📝 Abstract
While modern 3D reconstruction excels at modeling object geometry and appearance, it largely ignores the rich acoustic cues revealed through physical interaction. Object impact sounds convey material, stiffness, and structural properties that complement vision, yet existing impact sound modeling approaches either rely on expensive physics-based simulation or require large datasets to generalize in a purely data-driven manner. We introduce Audio-Visual Modal Sound Field (AV-MSF), a novel object-level acoustic representation reconstructed from multi-view images and only a few impact sound recordings. AV-MSF builds on 3D Gaussian Splatting integrated with dense 3D visual feature to provide a strong geometry-aware prior, and represents the impact sound field using compact, physically meaningful modal parameters, enabling robust few-shot reconstruction. Experiments on two real-world datasets show that AV-MSF achieves state-of-the-art impact sound rendering, outperforming both physics-based and data-driven baselines. Furthermore, we demonstrate downstream applications enabled by our representation, including contact localization and object sound editing.