audio editing

Edits and assembles digital audio by selecting, trimming, arranging, and crossfading clips across one or more tracks and by applying processing such as equalization, compression, noise reduction, time-stretching, and pitch correction. Produces final mixes and output files by managing levels and dynamics, routing and automation, and converting to required formats and technical specifications.

audioediting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Recomposer: Event-roll-guided generative audio editing

Sep 05, 2025
DP
Daniel P. W. Ellis
🏛️ Google DeepMind

To address the challenge of fine-grained single-event audio editing in complex soundscapes where acoustic sources overlap temporally, this paper proposes a generative audio editing method operating at millisecond resolution. Methodologically, it introduces a “recomposition” paradigm that jointly leverages text instructions, event-level semantic categories, and millisecond-accurate temporal maps as multimodal editing guidance. The architecture employs a SoundStream-based encoder-decoder Transformer, trained on paired synthetic audio data, and incorporates event-aware convolutional transcription to generate precise temporal supervision. Experiments demonstrate high-fidelity performance across three editing operations—deletion, insertion, and enhancement—in realistic, cluttered acoustic environments. Ablation studies confirm that editing quality is significantly influenced by the specificity of the action type, semantic category, and temporal precision, thereby validating the necessity and effectiveness of multimodal, joint modeling for audio editing.

Editing overlapping sound sources in complex scenesIntegrating textual and graphical event descriptions for editingUsing generative models for audio inpainting and enhancement

LLM2Fx-Tools: Tool Calling For Music Post-Production

Dec 01, 2025
SD
Seungheon Doh
🏛️ KAIST | Sony AI | Sony Group Corporation

This work addresses the challenge of enabling large language models (LLMs) to comprehend raw audio and autonomously generate executable audio effect chains (Fx-chains) for music post-production. We propose the first multimodal tool-calling framework tailored for audio effect synthesis, integrating audio representations, structured tool interfaces, chain-of-thought (CoT) planning, and autoregressive sequence modeling to achieve end-to-end mapping from input audio to effect types, ordering, and parameters. We introduce LP-Fx, a high-quality, human-annotated dataset for audio effect chaining, and pioneer the application of LLM tool-calling paradigms to audio processing. Experiments demonstrate that our system generates semantically coherent and parameter-plausible Fx-chains; successfully transfers processing characteristics in style-transfer tasks; and achieves strong interpretability and response fidelity, as validated by both human and LLM-based evaluation.

Generates executable audio effect sequences for music post-productionInfers effect chains from unprocessed and processed audio pairsTransfers audio effect styles from reference to new content

MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model

Dec 23, 2025
YT
Ye Tao
🏛️ Shanghai Jiao Tong University | Shanghai AI Laboratory | Nanjing University

Existing text-guided audio editing methods face three key limitations: (1) training-free approaches suffer from audio quality degradation due to diffusion inversion; (2) supervised methods are constrained by the scarcity of high-quality paired data and limited edit types; and (3) modality-decoupled architectures struggle to achieve fine-grained alignment between natural language instructions and acoustic features. This work first formally defines a comprehensive audio editing benchmark covering five fundamental operations—addition, replacement, deletion, reordering, and attribute modification—and introduces an event-level fine-grained annotation paradigm for synthetic data generation. We further propose a unified editing architecture featuring deep audio-language alignment: a Qwen2-Audio encoder, an MMDiT-based generator, and a custom joint instruction-tuning strategy. Experiments demonstrate 98.7% fidelity preservation in unedited regions, a 12.4% average improvement in editing accuracy over SOTA, and significant gains in spatial localization precision and instruction-following robustness.

Addresses limitations in text-guided audio editing methodsEnables precise cross-modal alignment for localized audio editingExtends task definitions to cover comprehensive audio editing operations

In real-world acoustic scenes, users struggle to manipulate unseparated mixed sound sources. To address this, we propose the first end-to-end, text-driven framework for real-time sound field editing, enabling direct, joint manipulation of multiple concurrent sources—such as “reduce air-conditioner noise and enhance speech”—via natural language instructions, without explicit source separation. Our method integrates large language model–based semantic parsing with a differentiable spectrogram decomposition–filtering–reconstruction architecture, supporting open-vocabulary, zero-shot editing. It is trained on a newly curated dataset comprising 160 hours of audio and 100,000 audio–text pairs. Experiments demonstrate significant improvements in source extraction, suppression, and level control: +3.2 dB in SI-SNR and +0.11 in STOI, while maintaining strong robustness and generalization across complex mixtures with 2–5 overlapping sources.

Control sound sources in mixtures via text instructionsEnhance auditory experience with semantic text filtersRemix multiple sounds simultaneously without separation

Latest Papers

What's happening recently
View more

This work proposes a fully hardware-based single-tone audio synthesis method implemented on an FPGA to meet the demanding requirements of professional audio applications for high-precision, low-latency sinusoidal signals. By leveraging digital signal synthesis techniques and optimized digital logic design, the approach efficiently generates highly stable sine waves at specific frequencies directly in hardware and integrates a digital-to-analog conversion interface for real-world audio output. Entirely eschewing conventional software or hybrid implementations, the solution achieves significantly reduced latency and enhanced timing accuracy through pure hardware execution. Experimental results demonstrate that the resulting synthesizer is well-suited for stringent electronic systems requiring exceptional signal fidelity—such as clock synchronization, communication transmission, and embedded control—offering high precision, minimal resource utilization, and strong real-time performance.

audio synthesizerdigital synthesisFPGA

Existing automatic music mixing approaches struggle to simultaneously achieve high-quality output and flexible style control. This work proposes Diff2Mix, the first end-to-end mixing system that integrates diffusion-based generative modeling with a differentiable audio mixer. The method enables global mix style guidance through reference audio while allowing users to explicitly adjust audio effect parameters, thereby offering a highly controllable mixing process. Experimental results demonstrate that Diff2Mix achieves state-of-the-art performance in both objective metrics and subjective listening tests, effectively balancing audio quality with editing flexibility.

automatic music mixingcontrollable generationdifferentiable audio effects

Existing audio generation evaluation methods struggle to simultaneously address the industrial sound design requirements of reference guidance, controllable variation, perceptual consistency, and workflow efficiency. This work proposes the first production-oriented evaluation framework for sound effect generation, structured around nine core production criteria and a two-stage protocol that enables systematic, goal-aligned comparison of heterogeneous generation and editing approaches. The framework integrates objective metrics—including Fréchet Audio Distance (FAD), ImageBind-based reference alignment, and diversity scores—with human listening experiments to holistically assess perceptual identity preservation and transient fidelity. Empirical results reveal complementary strengths across baseline methods, with AudioX achieving the best trade-off between reference alignment and output diversity while effectively supporting sound morphing tasks.

audio variationindustrial audio designproduction-oriented evaluation

This work addresses the challenge of imprecise control in existing speech editing methods under natural language instructions, which often suffer from semantic ambiguity in specifying edit types, parameters, and target regions. To overcome this, the authors propose a structured editing interface grounded in transcribed text, employing XML-style tags to explicitly denote operation types and anchor them to specific transcript spans or boundaries, thereby constructing a semantic timeline that circumvents the need for explicit time alignment. Building upon this framework, they enhance the continuous autoregressive model dots.tts to support four composable editing dimensions—lexical content, emotion, prosody, and pauses—while preserving contextual integrity in unedited segments. The contributions include the first structured instruction framework for speech editing, a task-oriented data curation pipeline, and doteBench, the first bilingual benchmark for precise evaluation. Experiments demonstrate state-of-the-art instruction-following accuracy and local fidelity across five editing tasks in doteBench, with audio quality comparable to leading open-source systems and no significant degradation in zero-shot TTS error rates or speaker similarity relative to the base model.

ambiguitycontent creationedit instruction

Hot Scholars

SL

Sichen Liu

MS Student, Huazhong University of Science and Technology
Generative ModelImage Generation
ZY

Zitong Yu

U.S. Food and Drug Administration
Medical imagingDeep learningMachine learningImage reconstruction
AS

Annelien Smets

imec-SMIT, Vrije Universiteit Brussel
recommender systemsdiscoveryserendipitymedia economics
ME

Micha Elsner

Assistant Professor of Linguistics, The Ohio State University
computational linguisticsBayesian methodsdiscourse structurelanguage acquisition