Relational Synthesis: Structure-Mediated Concatenative Synthesis for Foley and Retrieval-Augmented Audio Generation

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent tension between acoustic consistency and originality in Foley audio generation, alongside the information leakage problem in neural retrieval-augmented generation (RAG). To this end, we propose a relational synthesis method that pioneers the use of Gromov-Wasserstein structural costs in place of conventional concatenative objective costs. Rather than imitating content, our approach transfers the temporal dynamics of reference audio to reorganize source audio grains, thereby preserving acoustic characteristics while preventing direct copying. Experimental results demonstrate that the proposed method achieves superior performance in temporal coherence, acoustic fidelity, and leakage prevention metrics. Furthermore, it effectively maintains distribution-level quality and text-alignment capabilities.
📝 Abstract
We ask: given a retrieved source audio $S$ and a separate reference audio $R$, can we synthesize novel audio $Y$ out of this pair $(S,R)$ such that $Y$ remains acoustically consistent with $S$, while not persistently copying segments of $S$ or $R$? The first clause is a well-known goal in Foley audio production, and the second is a well-known issue in neural RAG when $S$ and $R$ are naively injected into neural generators. We show that both clauses can be addressed simultaneously using a method we coin relational synthesis, a variation of concatenative synthesis where target cost is replaced by a relational Gromov-like structural cost. Rather than imitating the content of $R$, relational synthesis exploits it from the"other side of the hill": it transfers the temporal structure and directed amplitude motion of $R$ to reorganize and concatenate the grains of $S$ in a novel manner that protects $S$'s acoustic information. Our experiments show that relational synthesis integrates naturally with neural RAG and produces Foley audio that performs well on metrics measuring temporal agreement, acoustic fidelity, and leakage persistence, while maintaining distribution-level quality and text alignment.
Problem

Research questions and friction points this paper is trying to address.

Foley audio generation
Retrieval-Augmented Generation
concatenative synthesis
audio leakage
acoustic consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Relational Synthesis
Concatenative Synthesis
Retrieval-Augmented Generation
Gromov-like Structural Cost
Foley Audio
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Keren Shao
University of California San Diego, La Jolla, CA, USA
A
Ayaka Kawano
University of California San Diego, La Jolla, CA, USA
Shlomo Dubnov
Shlomo Dubnov
Professor Music, Computer Science and Engineering, UCSD
Computer MusicMachine LearningComputational Creativity