🤖 AI Summary
This study addresses the limitation that outputs of existing audio generation models are non-editable, making it difficult to directly adjust notes, timbres, and modulation parameters. This work proposes unifying synthesizer programs into a sequential representation and employing autoregressive models to predict editable parameters and modulation routing from audio or text inputs. Methodologically, without requiring paired annotations or differentiable synthesizers, a two-stage training paradigm combining supervised learning with Group Relative Policy Optimization (GRPO) enables a single model to simultaneously support audio inversion and text-driven generation. To our knowledge, this is the first work to generate complete and fully editable synthesizer programs, achieving highly competitive performance on both tasks.
📝 Abstract
Audio generation models can translate natural-language descriptions into sound, but their outputs are typically waveforms. Their audio quality is constrained by audio compression, and their outputs do not readily support direct edits to notes, timbral parameters, or modulation relationships. We present AutoSynth, which represents MIDI performance events, fixed synthesizer parameters, and variable-length modulation routes as a unified sequence for a synthesizer, and learns their dependencies with an audio-conditioned autoregressive model. A single model supports both tasks. Given reference audio, the model directly predicts a synthesizer program; given text, it uses a pretrained audio generation model and converts the generated audio into a program. Training consists of two stages: supervised learning on large-scale audio-program pairs automatically constructed from a small set of native presets, followed by group-relative policy optimization with a mixed reward combining semantic similarity, pitch-related features, acoustic similarity, and sound usefulness. The pipeline requires neither paired text-target-program annotations nor a differentiable synthesizer. Experiments show that AutoSynth produces complete, editable synthesizer programs and achieves competitive results in both synthesizer inversion and text-driven generation. Audio demos and source code are available at https://auto-synth.github.io/.