InstructFX2FX: A Multi-turn Text-to-Preset Demo for Iterative Audio Effect Refinement

📅 2026-06-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of existing text-to-audio preset systems, which typically rely on one-shot mappings and thus struggle to support audio engineers in iteratively refining effect chains through multi-turn natural language instructions. To overcome this, we propose the first state-aware, interactive framework for multi-round audio effect tuning. Our approach leverages a large language model (LLM) as a high-level planner for effect selection and initial parameterization, coupled with a CLAP-guided perceptual optimization algorithm that enables progressive, state-preserving fine-tuning. Experiments on the SocialFX dataset demonstrate that, compared to a pure LLM re-prompting baseline, our method significantly reduces the MMD distance of target-oriented DSP features in 9 out of 10 descriptor transfer tasks, achieving an average reduction of 24% (from 0.45 to 0.34), thereby validating its superior controllability and stability.
📝 Abstract
We present InstructFX2FX, an interactive demonstration of multi-turn, text-guided audio effect refinement. Existing text-to-preset systems are largely single-shot, mapping one textual descriptor to one preset. Real audio engineering is instead sequential: engineers refine an existing effect chain through successive instructions such as "make it warmer" or "too harsh, soften it." This poses a stateful problem that single-shot systems do not address: given the current effect parameters and a new instruction, update the sound while preserving what earlier instructions already achieved. InstructFX2FX tackles this with a hybrid architecture that divides labor between a language model and CLAP-guided optimization. The LLM acts as a high-level planner that selects effects, orders the signal chain, and proposes initial parameters, motivated by recent evidence that LLM initialization outperforms CLAP optimization from a random start; CLAP-guided optimization then refines those parameters perceptually, giving more controllable updates than re-prompting the LLM at each turn, which tends to perturb settings the user did not ask to change. At the demo, attendees drive a dry recording through successive natural-language instructions and inspect the evolving FX chain, intermediate optimization checkpoints, and session state. On SocialFX-derived descriptor transitions, CLAP-guided refinement lowers target-directed DSP-feature MMD on 9 of 10 directed descriptor pairs relative to an LLM-only reprompting baseline, reducing the average MMD by roughly 24% (0.45 to 0.34).
Problem

Research questions and friction points this paper is trying to address.

text-to-preset
audio effect refinement
multi-turn interaction
stateful audio processing
iterative optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-turn interaction
text-to-preset
CLAP-guided optimization
audio effect refinement
hybrid architecture
S
Song-Ze Yu
Center for New Music and Audio Technologies (CNMAT), University of California, Berkeley
M
Milan Liessens Dujardin
Center for New Music and Audio Technologies (CNMAT), University of California, Berkeley
Y
Yuxuan Cai
Center for New Music and Audio Technologies (CNMAT), University of California, Berkeley
W
Wantong Zhang
Center for New Music and Audio Technologies (CNMAT), University of California, Berkeley