🤖 AI Summary
This work addresses a critical gap in audio-visual generation: while existing models produce semantically coherent and visually synchronized audio, they often violate fundamental acoustic physical laws and lack effective diagnostic tools. To bridge this gap, we propose the first diagnostic benchmark specifically designed for acoustic physical realism, formalizing eight measurable acoustic dimensions grounded in the sound generation, propagation, and reception pipeline. We further introduce a large-scale RGB-D audio-visual dataset with detailed acoustic annotations. Leveraging this benchmark, we develop a specialized prompt set and evaluator that integrate acoustic constraints into model training, reward modeling, and candidate selection. Experiments reveal significant deficiencies in mainstream models regarding basic acoustic processes, while our approach effectively guides optimization and substantially enhances the physical fidelity of generated audio.
📝 Abstract
Recent audio-video generators can produce semantically plausible and apparently synchronized sound, yet may still violate the acoustic processes implied by visible events and environments. Existing benchmarks provide limited support for attributing such violations to particular acoustic processes and quantifying their severity. We introduce AcoustiTrace, a diagnostic benchmark that formalizes acoustic physical realism in audio-video generation. AcoustiTrace organizes text-to-audio-video (T2AV) and image-to-audio-video (I2AV) evaluation around the acoustic process, covering sound generation, propagation environment, and acoustic reception through eight dimensions grounded in measurable acoustic quantities. Based on these evaluation dimensions, we construct a large-scale dataset organized around acoustic mechanisms, comprising real-world audio-video recordings and acoustically annotated RGB-D observations, and use it to develop targeted prompt suites and validated evaluators. Experiments reveal that even leading generators still struggle with fundamental acoustic processes despite producing plausible sound events. Finally, we show that the diagnostics AcoustiTrace provides for specific acoustic relations can guide model refinement toward more physically faithful audio and open new directions for incorporating acoustic principles into training objectives, reward modeling, and candidate selection.