Score
Designs and builds algorithms, tools, and workflows to edit and manipulate audio scenes and individual tracks, including selective insertion, deletion, and transformation of sound sources while preserving attributes such as speaker identity. Analyzes and implements audio quality control and artifact management to measure, constrain, and remediate degradations introduced by edits or synthesis.
In real-world acoustic scenes, users struggle to manipulate unseparated mixed sound sources. To address this, we propose the first end-to-end, text-driven framework for real-time sound field editing, enabling direct, joint manipulation of multiple concurrent sources—such as “reduce air-conditioner noise and enhance speech”—via natural language instructions, without explicit source separation. Our method integrates large language model–based semantic parsing with a differentiable spectrogram decomposition–filtering–reconstruction architecture, supporting open-vocabulary, zero-shot editing. It is trained on a newly curated dataset comprising 160 hours of audio and 100,000 audio–text pairs. Experiments demonstrate significant improvements in source extraction, suppression, and level control: +3.2 dB in SI-SNR and +0.11 in STOI, while maintaining strong robustness and generalization across complex mixtures with 2–5 overlapping sources.
To address the challenge of fine-grained single-event audio editing in complex soundscapes where acoustic sources overlap temporally, this paper proposes a generative audio editing method operating at millisecond resolution. Methodologically, it introduces a “recomposition” paradigm that jointly leverages text instructions, event-level semantic categories, and millisecond-accurate temporal maps as multimodal editing guidance. The architecture employs a SoundStream-based encoder-decoder Transformer, trained on paired synthetic audio data, and incorporates event-aware convolutional transcription to generate precise temporal supervision. Experiments demonstrate high-fidelity performance across three editing operations—deletion, insertion, and enhancement—in realistic, cluttered acoustic environments. Ablation studies confirm that editing quality is significantly influenced by the specificity of the action type, semantic category, and temporal precision, thereby validating the necessity and effectiveness of multimodal, joint modeling for audio editing.
Audio editing has long suffered from the lack of high-quality evaluation benchmarks and reliable automated assessment metrics. To address this, we propose an expert-knowledge-driven closed-loop evaluation framework. First, we construct AuditScore—the first subjective evaluation dataset for audio editing—comprising over 6,300 samples annotated with multi-dimensional professional ratings. Second, we train AuditEval, an automatic Mean Opinion Score (MOS) prediction model achieving high accuracy in quality estimation. Third, we leverage AuditEval in a reverse pipeline to filter and refine synthetic data, generating a validated pseudo-parallel dataset of superior quality. This work pioneers the organic integration of expert scoring, automated evaluation, and data curation: it introduces the first task-specific audio editing evaluation model and benchmark dataset, and establishes a “evaluate–feedback–generate” closed-loop paradigm—providing a reproducible, scalable foundation for future research and development in audio editing.
This study addresses a significant misalignment between current AI tools and the needs of practitioners in high-level narrative sound design. Employing a mixed-methods approach—combining 76 survey responses with semi-structured interviews of 20 industry professionals—the research systematically examines the current applications, challenges, and expectations surrounding AI in sound design workflows. From the perspective of sound designers, the study identifies five core themes for AI integration: context, workflow, potential, risks, and appropriate use, offering concrete design recommendations for developers. Findings indicate that while existing AI tools are effective for rapid, disposable audio tasks, they lack the narrative depth required for cinematic and immersive experiences. Practitioners express a clear preference for task-oriented, assistive AI systems over end-to-end generative solutions.
This work addresses the challenge of enabling large language models (LLMs) to comprehend raw audio and autonomously generate executable audio effect chains (Fx-chains) for music post-production. We propose the first multimodal tool-calling framework tailored for audio effect synthesis, integrating audio representations, structured tool interfaces, chain-of-thought (CoT) planning, and autoregressive sequence modeling to achieve end-to-end mapping from input audio to effect types, ordering, and parameters. We introduce LP-Fx, a high-quality, human-annotated dataset for audio effect chaining, and pioneer the application of LLM tool-calling paradigms to audio processing. Experiments demonstrate that our system generates semantically coherent and parameter-plausible Fx-chains; successfully transfers processing characteristics in style-transfer tasks; and achieves strong interpretability and response fidelity, as validated by both human and LLM-based evaluation.
This work addresses the challenge of imprecise control in existing speech editing methods under natural language instructions, which often suffer from semantic ambiguity in specifying edit types, parameters, and target regions. To overcome this, the authors propose a structured editing interface grounded in transcribed text, employing XML-style tags to explicitly denote operation types and anchor them to specific transcript spans or boundaries, thereby constructing a semantic timeline that circumvents the need for explicit time alignment. Building upon this framework, they enhance the continuous autoregressive model dots.tts to support four composable editing dimensions—lexical content, emotion, prosody, and pauses—while preserving contextual integrity in unedited segments. The contributions include the first structured instruction framework for speech editing, a task-oriented data curation pipeline, and doteBench, the first bilingual benchmark for precise evaluation. Experiments demonstrate state-of-the-art instruction-following accuracy and local fidelity across five editing tasks in doteBench, with audio quality comparable to leading open-source systems and no significant degradation in zero-shot TTS error rates or speaker similarity relative to the base model.
Current text-to-audio models implicitly handle sound scene construction, lacking explicit control and interpretability over the composition of sound events, their temporal arrangement, and the rendering process. This work proposes the first agent-based soundscape synthesis framework, which leverages a large language model to translate user intent into an editable scene plan. By integrating audio retrieval with on-demand generation, the framework enables controllable multi-event mixed rendering and supports human-in-the-loop interaction and tool selection. It explicitly models the full pipeline of sound scene creation—planning, source selection, layout, and rendering—producing audio that matches state-of-the-art text-to-audio models in both subjective listening quality and objective metrics. Moreover, downstream audio reasoning models trained on its synthetic data significantly outperform baselines trained solely on real-world recordings.