🤖 AI Summary
This study addresses the temporal-semantic inconsistencies arising from the isolated processing of audio and video in existing models by proposing a unified probabilistic framework that formulates joint generation, cross-modal generation, and editing as three subproblems under a single distribution. Building upon this framework, we construct a five-axis taxonomy that systematically organizes joint audio-visual editing tasks for the first time, mapping nine editing paradigms encompassing twenty-eight distinct types. By integrating multimodal fusion techniques with standardized evaluation protocols, this work identifies optimal methods, datasets, and metrics for various scenarios. Furthermore, it distills the most impactful open problems in the field, providing a comprehensive roadmap for future research in unified audio-visual editing.
📝 Abstract
Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the output kept coherent across modalities in time and semantics? A unified formulation casts joint generation, cross-modal generation, and joint editing as three problems defined on a single distribution over audio-visual pairs, and a taxonomy compares methods along five design axes. To our knowledge, this is the first overview to systematically taxonomize joint audio-visual editing, which we map as nine edit categories spanning 28 edit types. We describe methods, datasets, and metrics for each setting and close with the open problems we view as most consequential.