Score
Edits and assembles digital audio by selecting, trimming, arranging, and crossfading clips across one or more tracks and by applying processing such as equalization, compression, noise reduction, time-stretching, and pitch correction. Produces final mixes and output files by managing levels and dynamics, routing and automation, and converting to required formats and technical specifications.
To address the challenge of fine-grained single-event audio editing in complex soundscapes where acoustic sources overlap temporally, this paper proposes a generative audio editing method operating at millisecond resolution. Methodologically, it introduces a “recomposition” paradigm that jointly leverages text instructions, event-level semantic categories, and millisecond-accurate temporal maps as multimodal editing guidance. The architecture employs a SoundStream-based encoder-decoder Transformer, trained on paired synthetic audio data, and incorporates event-aware convolutional transcription to generate precise temporal supervision. Experiments demonstrate high-fidelity performance across three editing operations—deletion, insertion, and enhancement—in realistic, cluttered acoustic environments. Ablation studies confirm that editing quality is significantly influenced by the specificity of the action type, semantic category, and temporal precision, thereby validating the necessity and effectiveness of multimodal, joint modeling for audio editing.
This work addresses the challenge of enabling large language models (LLMs) to comprehend raw audio and autonomously generate executable audio effect chains (Fx-chains) for music post-production. We propose the first multimodal tool-calling framework tailored for audio effect synthesis, integrating audio representations, structured tool interfaces, chain-of-thought (CoT) planning, and autoregressive sequence modeling to achieve end-to-end mapping from input audio to effect types, ordering, and parameters. We introduce LP-Fx, a high-quality, human-annotated dataset for audio effect chaining, and pioneer the application of LLM tool-calling paradigms to audio processing. Experiments demonstrate that our system generates semantically coherent and parameter-plausible Fx-chains; successfully transfers processing characteristics in style-transfer tasks; and achieves strong interpretability and response fidelity, as validated by both human and LLM-based evaluation.
Existing text-guided audio editing methods face three key limitations: (1) training-free approaches suffer from audio quality degradation due to diffusion inversion; (2) supervised methods are constrained by the scarcity of high-quality paired data and limited edit types; and (3) modality-decoupled architectures struggle to achieve fine-grained alignment between natural language instructions and acoustic features. This work first formally defines a comprehensive audio editing benchmark covering five fundamental operations—addition, replacement, deletion, reordering, and attribute modification—and introduces an event-level fine-grained annotation paradigm for synthetic data generation. We further propose a unified editing architecture featuring deep audio-language alignment: a Qwen2-Audio encoder, an MMDiT-based generator, and a custom joint instruction-tuning strategy. Experiments demonstrate 98.7% fidelity preservation in unedited regions, a 12.4% average improvement in editing accuracy over SOTA, and significant gains in spatial localization precision and instruction-following robustness.
In real-world acoustic scenes, users struggle to manipulate unseparated mixed sound sources. To address this, we propose the first end-to-end, text-driven framework for real-time sound field editing, enabling direct, joint manipulation of multiple concurrent sources—such as “reduce air-conditioner noise and enhance speech”—via natural language instructions, without explicit source separation. Our method integrates large language model–based semantic parsing with a differentiable spectrogram decomposition–filtering–reconstruction architecture, supporting open-vocabulary, zero-shot editing. It is trained on a newly curated dataset comprising 160 hours of audio and 100,000 audio–text pairs. Experiments demonstrate significant improvements in source extraction, suppression, and level control: +3.2 dB in SI-SNR and +0.11 in STOI, while maintaining strong robustness and generalization across complex mixtures with 2–5 overlapping sources.
This work proposes a fully hardware-based single-tone audio synthesis method implemented on an FPGA to meet the demanding requirements of professional audio applications for high-precision, low-latency sinusoidal signals. By leveraging digital signal synthesis techniques and optimized digital logic design, the approach efficiently generates highly stable sine waves at specific frequencies directly in hardware and integrates a digital-to-analog conversion interface for real-world audio output. Entirely eschewing conventional software or hybrid implementations, the solution achieves significantly reduced latency and enhanced timing accuracy through pure hardware execution. Experimental results demonstrate that the resulting synthesizer is well-suited for stringent electronic systems requiring exceptional signal fidelity—such as clock synchronization, communication transmission, and embedded control—offering high precision, minimal resource utilization, and strong real-time performance.
Existing automatic music mixing approaches struggle to simultaneously achieve high-quality output and flexible style control. This work proposes Diff2Mix, the first end-to-end mixing system that integrates diffusion-based generative modeling with a differentiable audio mixer. The method enables global mix style guidance through reference audio while allowing users to explicitly adjust audio effect parameters, thereby offering a highly controllable mixing process. Experimental results demonstrate that Diff2Mix achieves state-of-the-art performance in both objective metrics and subjective listening tests, effectively balancing audio quality with editing flexibility.
Existing audio generation evaluation methods struggle to simultaneously address the industrial sound design requirements of reference guidance, controllable variation, perceptual consistency, and workflow efficiency. This work proposes the first production-oriented evaluation framework for sound effect generation, structured around nine core production criteria and a two-stage protocol that enables systematic, goal-aligned comparison of heterogeneous generation and editing approaches. The framework integrates objective metrics—including Fréchet Audio Distance (FAD), ImageBind-based reference alignment, and diversity scores—with human listening experiments to holistically assess perceptual identity preservation and transient fidelity. Empirical results reveal complementary strengths across baseline methods, with AudioX achieving the best trade-off between reference alignment and output diversity while effectively supporting sound morphing tasks.
This work addresses the challenge of imprecise control in existing speech editing methods under natural language instructions, which often suffer from semantic ambiguity in specifying edit types, parameters, and target regions. To overcome this, the authors propose a structured editing interface grounded in transcribed text, employing XML-style tags to explicitly denote operation types and anchor them to specific transcript spans or boundaries, thereby constructing a semantic timeline that circumvents the need for explicit time alignment. Building upon this framework, they enhance the continuous autoregressive model dots.tts to support four composable editing dimensions—lexical content, emotion, prosody, and pauses—while preserving contextual integrity in unedited segments. The contributions include the first structured instruction framework for speech editing, a task-oriented data curation pipeline, and doteBench, the first bilingual benchmark for precise evaluation. Experiments demonstrate state-of-the-art instruction-following accuracy and local fidelity across five editing tasks in doteBench, with audio quality comparable to leading open-source systems and no significant degradation in zero-shot TTS error rates or speaker similarity relative to the base model.