Spot, Separate, and Enhance: Fully Generative Approach for Audio Mixing

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the tasks of video audio remixing, denoising, and dereverberation by proposing SSE, the first multimodal user-guided generative audio processing model. By integrating visual cues with textual instructions, SSE enables natural language-driven joint audio optimization. The core contributions include the construction of the first fully generative audio mixing framework, the introduction of the DegradedMix benchmark dataset, and the proposal of novel generative evaluation metrics. Experimental results demonstrate that SSE significantly outperforms existing baseline methods in terms of both controllability and generation quality.
📝 Abstract
We introduce Spot, Separate, and Enhance (SSE), the first multimodal, user-guided generative model for audio remixing and enhancement. SSE enhances video content by rebalancing the audio, removing unwanted audio sources, and reducing reverberation, guided by both video and textual descriptions. To support its training and evaluation, we propose DegradedMix, a new dataset built on the audio remixing benchmark MuddyMix. We also adopt evaluation metrics from generative modeling, which better capture the creative nature of remixing than standard reconstruction-based metrics. SSE outperforms existing baselines in both controllability and remixing quality, as shown by extensive experiments. Project page: https://sse-ai.notion.site
Problem

Research questions and friction points this paper is trying to address.

audio remixing
audio enhancement
multimodal generation
source separation
reverberation reduction
Innovation

Methods, ideas, or system contributions that make the work stand out.

generative audio remixing
multimodal guidance
audio enhancement
DegradedMix dataset
user-guided generation
🔎 Similar Papers