audio scene editing

Designs and builds algorithms, tools, and workflows to edit and manipulate audio scenes and individual tracks, including selective insertion, deletion, and transformation of sound sources while preserving attributes such as speaker identity. Analyzes and implements audio quality control and artifact management to measure, constrain, and remediate degradations introduced by edits or synthesis.

audiosceneediting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

In real-world acoustic scenes, users struggle to manipulate unseparated mixed sound sources. To address this, we propose the first end-to-end, text-driven framework for real-time sound field editing, enabling direct, joint manipulation of multiple concurrent sources—such as “reduce air-conditioner noise and enhance speech”—via natural language instructions, without explicit source separation. Our method integrates large language model–based semantic parsing with a differentiable spectrogram decomposition–filtering–reconstruction architecture, supporting open-vocabulary, zero-shot editing. It is trained on a newly curated dataset comprising 160 hours of audio and 100,000 audio–text pairs. Experiments demonstrate significant improvements in source extraction, suppression, and level control: +3.2 dB in SI-SNR and +0.11 in STOI, while maintaining strong robustness and generalization across complex mixtures with 2–5 overlapping sources.

Control sound sources in mixtures via text instructionsEnhance auditory experience with semantic text filtersRemix multiple sounds simultaneously without separation

Recomposer: Event-roll-guided generative audio editing

Sep 05, 2025
DP
Daniel P. W. Ellis
🏛️ Google DeepMind

To address the challenge of fine-grained single-event audio editing in complex soundscapes where acoustic sources overlap temporally, this paper proposes a generative audio editing method operating at millisecond resolution. Methodologically, it introduces a “recomposition” paradigm that jointly leverages text instructions, event-level semantic categories, and millisecond-accurate temporal maps as multimodal editing guidance. The architecture employs a SoundStream-based encoder-decoder Transformer, trained on paired synthetic audio data, and incorporates event-aware convolutional transcription to generate precise temporal supervision. Experiments demonstrate high-fidelity performance across three editing operations—deletion, insertion, and enhancement—in realistic, cluttered acoustic environments. Ablation studies confirm that editing quality is significantly influenced by the specificity of the action type, semantic category, and temporal precision, thereby validating the necessity and effectiveness of multimodal, joint modeling for audio editing.

Editing overlapping sound sources in complex scenesIntegrating textual and graphical event descriptions for editingUsing generative models for audio inpainting and enhancement

Audio editing has long suffered from the lack of high-quality evaluation benchmarks and reliable automated assessment metrics. To address this, we propose an expert-knowledge-driven closed-loop evaluation framework. First, we construct AuditScore—the first subjective evaluation dataset for audio editing—comprising over 6,300 samples annotated with multi-dimensional professional ratings. Second, we train AuditEval, an automatic Mean Opinion Score (MOS) prediction model achieving high accuracy in quality estimation. Third, we leverage AuditEval in a reverse pipeline to filter and refine synthetic data, generating a validated pseudo-parallel dataset of superior quality. This work pioneers the organic integration of expert scoring, automated evaluation, and data curation: it introduces the first task-specific audio editing evaluation model and benchmark dataset, and establishes a “evaluate–feedback–generate” closed-loop paradigm—providing a reproducible, scalable foundation for future research and development in audio editing.

Absence of comprehensive evaluation metrics for audio editing qualityLack of high-quality benchmark datasets for audio editing tasksNeed for expert-informed methods to construct pseudo-parallel datasets

This study addresses a significant misalignment between current AI tools and the needs of practitioners in high-level narrative sound design. Employing a mixed-methods approach—combining 76 survey responses with semi-structured interviews of 20 industry professionals—the research systematically examines the current applications, challenges, and expectations surrounding AI in sound design workflows. From the perspective of sound designers, the study identifies five core themes for AI integration: context, workflow, potential, risks, and appropriate use, offering concrete design recommendations for developers. Findings indicate that while existing AI tools are effective for rapid, disposable audio tasks, they lack the narrative depth required for cinematic and immersive experiences. Practitioners express a clear preference for task-oriented, assistive AI systems over end-to-end generative solutions.

AI integrationcreative workflowshuman-AI collaboration

LLM2Fx-Tools: Tool Calling For Music Post-Production

Dec 01, 2025
SD
Seungheon Doh
🏛️ KAIST | Sony AI | Sony Group Corporation

This work addresses the challenge of enabling large language models (LLMs) to comprehend raw audio and autonomously generate executable audio effect chains (Fx-chains) for music post-production. We propose the first multimodal tool-calling framework tailored for audio effect synthesis, integrating audio representations, structured tool interfaces, chain-of-thought (CoT) planning, and autoregressive sequence modeling to achieve end-to-end mapping from input audio to effect types, ordering, and parameters. We introduce LP-Fx, a high-quality, human-annotated dataset for audio effect chaining, and pioneer the application of LLM tool-calling paradigms to audio processing. Experiments demonstrate that our system generates semantically coherent and parameter-plausible Fx-chains; successfully transfers processing characteristics in style-transfer tasks; and achieves strong interpretability and response fidelity, as validated by both human and LLM-based evaluation.

Generates executable audio effect sequences for music post-productionInfers effect chains from unprocessed and processed audio pairsTransfers audio effect styles from reference to new content

Latest Papers

What's happening recently
View more

This work addresses the challenge of imprecise control in existing speech editing methods under natural language instructions, which often suffer from semantic ambiguity in specifying edit types, parameters, and target regions. To overcome this, the authors propose a structured editing interface grounded in transcribed text, employing XML-style tags to explicitly denote operation types and anchor them to specific transcript spans or boundaries, thereby constructing a semantic timeline that circumvents the need for explicit time alignment. Building upon this framework, they enhance the continuous autoregressive model dots.tts to support four composable editing dimensions—lexical content, emotion, prosody, and pauses—while preserving contextual integrity in unedited segments. The contributions include the first structured instruction framework for speech editing, a task-oriented data curation pipeline, and doteBench, the first bilingual benchmark for precise evaluation. Experiments demonstrate state-of-the-art instruction-following accuracy and local fidelity across five editing tasks in doteBench, with audio quality comparable to leading open-source systems and no significant degradation in zero-shot TTS error rates or speaker similarity relative to the base model.

ambiguitycontent creationedit instruction

Current text-to-audio models implicitly handle sound scene construction, lacking explicit control and interpretability over the composition of sound events, their temporal arrangement, and the rendering process. This work proposes the first agent-based soundscape synthesis framework, which leverages a large language model to translate user intent into an editable scene plan. By integrating audio retrieval with on-demand generation, the framework enables controllable multi-event mixed rendering and supports human-in-the-loop interaction and tool selection. It explicitly models the full pipeline of sound scene creation—planning, source selection, layout, and rendering—producing audio that matches state-of-the-art text-to-audio models in both subjective listening quality and objective metrics. Moreover, downstream audio reasoning models trained on its synthetic data significantly outperform baselines trained solely on real-world recordings.

audio-language supervisioncompositional audiocontrollable audio generation

Hot Scholars

ZW

Zhizheng Wu

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), Mel Lab
Spoken Language ProcessingDeepFake detectionMusic Processing
HH

Hossein Hassani

Computer Science and Engineering, University of Kurdistan Hewlêr, Kurdistan Region, Iraq.
Natural Language ProcessingComputational LinguisticsMachine Learning
WW

Wenwu Wang

Professor, University of Surrey, UK
signal processingmachine learningmachine listeningaudio/speech/audio-visual
GT

George Tzanetakis

Professor of Computer Science, Faculty of Engineering, University of Victoria
music information retrievalaudio signal processingmachine learninghuman-computer interaction
GM

Gallil Maimon

Hebrew University of Jerusalem
Machine LearningArtificial IntelligenceSpeech and Audio ProcessingNatural Language Processing