FillGauss: Fine-Grained Filling-Aware Impact Sound Generation for 3D Gaussian Splatting

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the physical implausibility in existing impact sound synthesis methods, which largely ignore how an object’s internal filling state—such as content type and fill level—affects acoustic resonance and damping. To bridge this gap, we introduce a fine-grained, fill-aware impact sound generation task and present FillImpact, a novel multimodal dataset comprising over 5,000 real-world audio recordings annotated with geometric, impact location, and filling condition metadata. We propose FillGauss, a physics-informed framework built upon 3D Gaussian Splatting and latent diffusion architectures, which jointly conditions on geometry, impact position, and filling state to generate high-fidelity, physically consistent sounds. Experiments demonstrate that our approach significantly outperforms prior methods in both physical plausibility and perceptual quality, establishing a new state of the art in physics-driven cross-modal audio generation.
📝 Abstract
Synthesizing physically plausible impact sounds from visual observations remains a great challenge in multi-modal AI. Existing 3D-aware audio generation methods primarily model the surface geometry of hollow rigid bodies. However, they fundamentally overlook internal filling states, a critical physical factor that drastically modulates acoustic resonance and damping. To address this issue, we have defined a new task called Fine-Grained Filling-Aware Impact Sound Generation. As a foundational step, we first introduce the fine-grained fill-aware dataset (FillImpact), a pioneering multi-modal collection comprising over 5,000 rigorous acoustic recordings from 88 diverse real-world objects. It captures impact interactions with varying internal contents (i.e., water, rice), a continuous range of fill levels, and distinct striker materials. Furthermore, comprehensive acoustic analysis confirms that the collected data closely aligns with established physical laws governing acoustic resonance and damping, indicating its suitability for physically grounded modeling. Building on this dataset, we propose a novel generative framework (FillGauss) that integrates 3D Gaussian Splatting (3DGS) with internal state conditioning for sound generation. By fusing 3DGS geometric features, precise 3D spatial strike coordinates, and fine-grained textual physical conditions within a latent diffusion architecture, FillGauss enables position-aware, striker-aware, and filling-aware audio generation. Extensive experiments demonstrate that our approach could generate high-fidelity impact sounds that adhere to underlying physical principles, establishing a new state-of-the-art for physically grounded cross-modal audio generation.
Problem

Research questions and friction points this paper is trying to address.

impact sound generation
filling-aware
3D Gaussian Splatting
physical plausibility
multi-modal AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Filling-Aware Sound Generation
3D Gaussian Splatting
Physical Audio Synthesis
Multimodal Dataset
Latent Diffusion Model
🔎 Similar Papers
No similar papers found.