Score
Designs and implements algorithms and pipelines that programmatically generate, synthesize, and modify audio signals and clips — including procedural sound events, codec-based renderings, synthetic clip generation, and augmentation techniques. Produces controlled temporal event placements, diverse event variations and durations, and mixed background contexts to cheaply scale labeled temporal training data and support model training and evaluation.
To address the challenge of fine-grained single-event audio editing in complex soundscapes where acoustic sources overlap temporally, this paper proposes a generative audio editing method operating at millisecond resolution. Methodologically, it introduces a “recomposition” paradigm that jointly leverages text instructions, event-level semantic categories, and millisecond-accurate temporal maps as multimodal editing guidance. The architecture employs a SoundStream-based encoder-decoder Transformer, trained on paired synthetic audio data, and incorporates event-aware convolutional transcription to generate precise temporal supervision. Experiments demonstrate high-fidelity performance across three editing operations—deletion, insertion, and enhancement—in realistic, cluttered acoustic environments. Ablation studies confirm that editing quality is significantly influenced by the specificity of the action type, semantic category, and temporal precision, thereby validating the necessity and effectiveness of multimodal, joint modeling for audio editing.
This work addresses the challenge of simultaneously modeling fine-grained event-level timing and long-range structural coherence in audio-driven music game level generation. To this end, we propose a multimodal sequence-to-sequence approach that conditions on both audio segments and level metadata, employing an event-based tokenized representation that explicitly encodes beat-aligned actions and their relative temporal offsets—thereby overcoming the limitations of conventional frame-level representations. Built upon a Transformer architecture, our model jointly learns the alignment between audio and event sequences to generate levels with precise rhythmic fidelity. Experimental results demonstrate that the proposed method significantly outperforms frame-level baselines on event-level evaluation metrics and enables systematic analysis of how audio cues contribute to rhythm-aligned prediction.
This work addresses the challenge of enabling large language models (LLMs) to comprehend raw audio and autonomously generate executable audio effect chains (Fx-chains) for music post-production. We propose the first multimodal tool-calling framework tailored for audio effect synthesis, integrating audio representations, structured tool interfaces, chain-of-thought (CoT) planning, and autoregressive sequence modeling to achieve end-to-end mapping from input audio to effect types, ordering, and parameters. We introduce LP-Fx, a high-quality, human-annotated dataset for audio effect chaining, and pioneer the application of LLM tool-calling paradigms to audio processing. Experiments demonstrate that our system generates semantically coherent and parameter-plausible Fx-chains; successfully transfers processing characteristics in style-transfer tasks; and achieves strong interpretability and response fidelity, as validated by both human and LLM-based evaluation.
This work addresses the problem of jointly conditioning high-fidelity audio generation on interpretable time-varying control signals (e.g., loudness, brightness, pitch) and text prompts, while supporting sketch-based control for vocal onomatopoeia or acoustic contours. Methodologically, we propose a lightweight linear control adapter—requiring only a single linear layer and 40k fine-tuning steps—and introduce stochastic median filtering to robustly model multi-granularity onomatopoeic timing. We unify text, time-varying signals, and sketch conditions within a latent diffusion Transformer (DiT) framework. Experiments demonstrate that our approach significantly improves fidelity to reference acoustic contours while preserving textual semantic consistency and audio quality—outperforming text-only baselines. The method establishes an efficient, artist-centric paradigm for fine-grained acoustic creation with precise, intuitive control.
To address temporal misalignment, semantic inconsistency, and reliance on costly manual timestamp annotations in video-driven Foley sound synthesis, this paper proposes Video2Sound—a fully end-to-end self-supervised framework. Methodologically, it introduces RMS intensity envelopes as unsupervised temporal event cues for the first time, enabling fine-grained control via RMS discretization and a novel RMS-ControlNet that jointly incorporates textual and audio-semantic prompts. The architecture adopts a two-stage pipeline—Video2RMS followed by RMS2Sound—integrating self-supervised learning, ControlNet-guided diffusion modeling, pretrained text-to-audio (T2A) priors, and explicit RMS feature representation. Extensive evaluation demonstrates state-of-the-art performance in audiovisual alignment, impact timing accuracy, dynamic intensity modeling, timbral fidelity, and controllability over sound details. The code, pretrained models, and interactive demo are publicly released.
This study addresses the growing demand in digital applications for high-quality, semantically aligned, and contextually relevant AI-generated sound effects that exhibit both diversity and controllability. It presents a systematic review of sound effect generation models from the past five years, encompassing approaches driven by text, visual, audio, and multimodal inputs. By analyzing 30 peer-reviewed papers sourced from Google Scholar, IEEE Xplore, and ACM Digital Library, this work offers the first comprehensive comparison of how different input modalities influence generation performance. The review highlights advances in audio fidelity, semantic alignment, and temporal coherence, while identifying persistent challenges such as temporal synchronization and perceptual consistency. Furthermore, it underscores a significant gap between current objective evaluation metrics and human perceptual judgments.
This study addresses the limitation of existing audio effects research, which predominantly focuses on parameter mapping while neglecting the construction of internal control spaces. We propose a novel design paradigm that transforms synthesizer parameter spaces into effect control spaces, restructuring 78 synthesizer parameters into timbre-shaping controls to enable reusable processing chain designs. Methodologically, we employ a retrieval-augmented A&R-CTAG model to generate text-conditioned configurations, validated through signal-level analysis. Experimental results demonstrate that the proposed role-based mapping approach yields significant spectral differences compared to random mapping baselines, effectively enhancing both the controllability and reusability of audio effects processors.
This study addresses the limitation that outputs of existing audio generation models are non-editable, making it difficult to directly adjust notes, timbres, and modulation parameters. This work proposes unifying synthesizer programs into a sequential representation and employing autoregressive models to predict editable parameters and modulation routing from audio or text inputs. Methodologically, without requiring paired annotations or differentiable synthesizers, a two-stage training paradigm combining supervised learning with Group Relative Policy Optimization (GRPO) enables a single model to simultaneously support audio inversion and text-driven generation. To our knowledge, this is the first work to generate complete and fully editable synthesizer programs, achieving highly competitive performance on both tasks.
This work addresses the limitation of existing text-to-audio preset systems, which typically rely on one-shot mappings and thus struggle to support audio engineers in iteratively refining effect chains through multi-turn natural language instructions. To overcome this, we propose the first state-aware, interactive framework for multi-round audio effect tuning. Our approach leverages a large language model (LLM) as a high-level planner for effect selection and initial parameterization, coupled with a CLAP-guided perceptual optimization algorithm that enables progressive, state-preserving fine-tuning. Experiments on the SocialFX dataset demonstrate that, compared to a pure LLM re-prompting baseline, our method significantly reduces the MMD distance of target-oriented DSP features in 9 out of 10 descriptor transfer tasks, achieving an average reduction of 24% (from 0.45 to 0.34), thereby validating its superior controllability and stability.
Existing audio generation models lack comprehensive evaluation across multiple dimensions—such as semantic fidelity, speaker consistency, and temporal control—in the context of mixed audio. To address this gap, this work introduces the first fine-grained, multi-control benchmark for mixed audio generation, comprising approximately 4,000 human-verified speech, music, and sound effect segments with precise temporal alignments, along with dedicated subsets for voice cloning and temporally conditioned generation. Leveraging multimodal human annotations and a multidimensional evaluation protocol—including acoustic fidelity, speech quality, semantic alignment, and temporal accuracy—we systematically assess state-of-the-art models. Our evaluation reveals significant trade-offs among these capabilities, with no single model achieving consistent superiority across all dimensions, thereby highlighting the core challenges in controllable mixed audio generation.