AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of scarce paired supervision data and insufficient cross-modal consistency in editing sounding objects within audio-visual scenes. To this end, it constructs AVIOBench, the first large-scale paired audio-visual editing benchmark, and proposes the AVIO model. By leveraging source-conditioned feature modulation, AVIO achieves joint audio-video editing while introducing a reference-frame curriculum learning mechanism to support instruction-based editing within a single unified model. The framework further integrates visual grounding, cross-modal consistency verification, and fine-tuning of pretrained generative models. Quantitative and qualitative evaluations demonstrate that the proposed method effectively performs object insertion and removal, supporting optional visual guidance for precise control over appearance and spatial positioning.
📝 Abstract
Adding or removing a sounding object requires coordinated changes to visual content and sound while preserving the surrounding scene. Yet paired supervision for localized non-speech audiovisual editing remains limited, as visual and acoustic edits must target the same object and isolate its sound from overlapping sources. To address this gap, we introduce \textit{AVIOBench}, a dataset comprising 37.9 hours of paired audiovisual examples spanning 1{,}878 target-object names. AVIOBench links the visual presence and acoustic contribution of each target object through a shared identity and visual mask. Our automated pipeline uses visual grounding and cross-modal consistency to select target-sound removal candidates, then jointly refines the audiovisual pairs to improve perceptual quality and cross-modal consistency. Building on this dataset, we propose \textit{AVIO}, which adapts a pretrained text-to audiovisual generation model through source-conditioned feature modulation to jointly learn object addition and removal. A reference-frame curriculum gradually reduces reference conditioning during training, enabling one model to perform instruction-only editing with optional visual guidance. Quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition, with optional reference guidance providing appearance and placement control for addition.
Problem

Research questions and friction points this paper is trying to address.

audiovisual editing
sounding object
object addition and removal
paired supervision
cross-modal consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audiovisual Editing
Source-Conditioned Feature Modulation
Reference-Frame Curriculum
Cross-Modal Consistency
AVIOBench
🔎 Similar Papers
No similar papers found.