Soundwich: Video Generation with Layered and Controllable Audio

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of source-level control in existing audio-visual generation models caused by audio mixing. It proposes a training-free framework that transforms a frozen flow matching model into a multi-track generator. Methodologically, the approach introduces a shared scene representation to maintain cross-modal temporal coherence and optimizes a cross-modal interaction routing mechanism corresponding to visual sources to achieve precise separation. This work significantly improves temporal control precision, source separability, and generation naturalness, enabling flexible independent editing operations such as track retiming, muting, and replacement.
📝 Abstract
Recent joint audio-video generative models can synthesize realistic videos with synchronized sound, but typically generate audio as a single mixed track. This limits source-level control and differs from practical audiovisual workflows, where speech, music, sound effects, and ambient sounds are represented as separate editable tracks. We introduce Soundwich, a training-free framework that transforms a frozen joint audio-video flow-matching model into a generator of multiple synchronized, independently editable audio stems coupled to a shared video. Soundwich generates separate audio stems with explicit control over their temporal activity. To keep separately generated sounds coherent, we introduce a shared scene representation that communicates global audiovisual context across stems while preserving their source-level separation. We further route cross-modal interactions between each audio stem and its corresponding visual source, improving audiovisual consistency. The resulting stems remain synchronized with the video and can be independently retimed, muted, replaced, or remixed. Experiments and human evaluations show improved temporal control, source separation, and naturalness, while enabling flexible source-level editing within coherent audiovisual generation. Code is available at https://github.com/CodyNing/Soundwich.
Problem

Research questions and friction points this paper is trying to address.

video generation
audio-visual synthesis
source separation
controllable audio
multi-track generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

training-free framework
audio stems generation
shared scene representation
cross-modal routing
controllable video generation
Z
Zhuo Ning
Simon Fraser University
A
AmirHossein Naghi Razlighi
University of Cyprus; CYENS Centre of Excellence
S
Sagi Polaczek
Tel Aviv University
Daniel Cohen-Or
Daniel Cohen-Or
Professor of Computer Science, Tel Aviv University
GraphicsImagingGeometric Modeling
A
Ali Mahdavi-Amiri
Simon Fraser University