Score
Designs, builds, and analyzes artifacts and systems for capture, synthesis, processing, transformation, and evaluation of audio (sound recordings, speech, musical signals) and their auditory elements. Work includes signal processing pipelines, feature extraction, mixing/encoding, noise reduction, and organization of auditory elements for human playback or machine consumption.
This work addresses the challenge of enabling large language models (LLMs) to comprehend raw audio and autonomously generate executable audio effect chains (Fx-chains) for music post-production. We propose the first multimodal tool-calling framework tailored for audio effect synthesis, integrating audio representations, structured tool interfaces, chain-of-thought (CoT) planning, and autoregressive sequence modeling to achieve end-to-end mapping from input audio to effect types, ordering, and parameters. We introduce LP-Fx, a high-quality, human-annotated dataset for audio effect chaining, and pioneer the application of LLM tool-calling paradigms to audio processing. Experiments demonstrate that our system generates semantically coherent and parameter-plausible Fx-chains; successfully transfers processing characteristics in style-transfer tasks; and achieves strong interpretability and response fidelity, as validated by both human and LLM-based evaluation.
This study addresses a significant misalignment between current AI tools and the needs of practitioners in high-level narrative sound design. Employing a mixed-methods approach—combining 76 survey responses with semi-structured interviews of 20 industry professionals—the research systematically examines the current applications, challenges, and expectations surrounding AI in sound design workflows. From the perspective of sound designers, the study identifies five core themes for AI integration: context, workflow, potential, risks, and appropriate use, offering concrete design recommendations for developers. Findings indicate that while existing AI tools are effective for rapid, disposable audio tasks, they lack the narrative depth required for cinematic and immersive experiences. Practitioners express a clear preference for task-oriented, assistive AI systems over end-to-end generative solutions.
This work addresses the black-box nature of music mixing by proposing a differentiable graph-structured inverse engineering method to automatically infer the processing chain and combination topology applied to dry source signals from the final mix. The approach constructs a parameterized mixing graph using differentiable audio processors and jointly optimizes both graph structure and parameters via gradient descent, augmented by a dry–wet ratio-guided iterative pruning strategy. Compared to manual modeling, it achieves high-fidelity mix reconstruction (PESQ ≥ 3.2, STOI ≥ 0.92) while removing approximately 67% of redundant processors, substantially reducing model complexity. The framework supports batch-parallel optimization and scalable generation of large mixing graphs. Experiments demonstrate its efficiency, scalability, and perceptual naturalness, establishing a novel paradigm for mixing analysis, AI-assisted production, and interpretable audio processing.
In real-world acoustic scenes, users struggle to manipulate unseparated mixed sound sources. To address this, we propose the first end-to-end, text-driven framework for real-time sound field editing, enabling direct, joint manipulation of multiple concurrent sources—such as “reduce air-conditioner noise and enhance speech”—via natural language instructions, without explicit source separation. Our method integrates large language model–based semantic parsing with a differentiable spectrogram decomposition–filtering–reconstruction architecture, supporting open-vocabulary, zero-shot editing. It is trained on a newly curated dataset comprising 160 hours of audio and 100,000 audio–text pairs. Experiments demonstrate significant improvements in source extraction, suppression, and level control: +3.2 dB in SI-SNR and +0.11 in STOI, while maintaining strong robustness and generalization across complex mixtures with 2–5 overlapping sources.
To address the longstanding reliance on single-task models and the lack of generalizable representations in computer audition, this paper proposes a systematic framework for constructing Auditory Foundation Models (AFMs). Methodologically, it establishes the core paradigm of AFMs for the first time, integrating unified multi-task modeling, cross-modal (audio–text) aligned representation learning, and instruction-driven human–machine interaction. Technically, the framework encompasses large-scale audio–text contrastive pretraining, multi-task prompt tuning, and self-supervised audio modeling. Experiments demonstrate that the proposed AFM achieves substantial performance gains across 10+ downstream tasks—including automatic speech recognition, sound source separation, and environmental sound classification—while enabling zero-shot transfer and open-domain speech understanding. This work advances computer audition toward generality, multi-task synergy, and natural human–machine interaction.
This work addresses the lack of a unified, high-quality transcoding method for spatial audio across diverse acquisition formats—such as Ambisonics or microphone arrays—and arbitrary playback systems. The authors propose a general parametric framework that estimates spatial metadata of primary sources and ambient sound in the time–frequency domain, constructs a spatial covariance model tailored to the target playback setup, and derives an optimal linear downmix matrix. This approach supports independent rotation between acquisition and playback geometries and, for the first time, unifies processing for both Ambisonics and raw microphone array inputs. It accommodates arbitrary array configurations, variable numbers of sources, and arbitrary angular power distributions of ambient sound. Listening tests demonstrate that the method significantly outperforms existing parametric renderers across various content types and playback configurations, with particularly notable perceptual improvements for low-order or geometrically constrained arrays.
Existing audio generation evaluation methods struggle to simultaneously address the industrial sound design requirements of reference guidance, controllable variation, perceptual consistency, and workflow efficiency. This work proposes the first production-oriented evaluation framework for sound effect generation, structured around nine core production criteria and a two-stage protocol that enables systematic, goal-aligned comparison of heterogeneous generation and editing approaches. The framework integrates objective metrics—including Fréchet Audio Distance (FAD), ImageBind-based reference alignment, and diversity scores—with human listening experiments to holistically assess perceptual identity preservation and transient fidelity. Empirical results reveal complementary strengths across baseline methods, with AudioX achieving the best trade-off between reference alignment and output diversity while effectively supporting sound morphing tasks.