audio

Designs, builds, and analyzes artifacts and systems for capture, synthesis, processing, transformation, and evaluation of audio (sound recordings, speech, musical signals) and their auditory elements. Work includes signal processing pipelines, feature extraction, mixing/encoding, noise reduction, and organization of auditory elements for human playback or machine consumption.

audio

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.95
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$214K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

LLM2Fx-Tools: Tool Calling For Music Post-Production

Dec 01, 2025
SD
Seungheon Doh
🏛️ KAIST | Sony AI | Sony Group Corporation

This work addresses the challenge of enabling large language models (LLMs) to comprehend raw audio and autonomously generate executable audio effect chains (Fx-chains) for music post-production. We propose the first multimodal tool-calling framework tailored for audio effect synthesis, integrating audio representations, structured tool interfaces, chain-of-thought (CoT) planning, and autoregressive sequence modeling to achieve end-to-end mapping from input audio to effect types, ordering, and parameters. We introduce LP-Fx, a high-quality, human-annotated dataset for audio effect chaining, and pioneer the application of LLM tool-calling paradigms to audio processing. Experiments demonstrate that our system generates semantically coherent and parameter-plausible Fx-chains; successfully transfers processing characteristics in style-transfer tasks; and achieves strong interpretability and response fidelity, as validated by both human and LLM-based evaluation.

Generates executable audio effect sequences for music post-productionInfers effect chains from unprocessed and processed audio pairsTransfers audio effect styles from reference to new content

This study addresses a significant misalignment between current AI tools and the needs of practitioners in high-level narrative sound design. Employing a mixed-methods approach—combining 76 survey responses with semi-structured interviews of 20 industry professionals—the research systematically examines the current applications, challenges, and expectations surrounding AI in sound design workflows. From the perspective of sound designers, the study identifies five core themes for AI integration: context, workflow, potential, risks, and appropriate use, offering concrete design recommendations for developers. Findings indicate that while existing AI tools are effective for rapid, disposable audio tasks, they lack the narrative depth required for cinematic and immersive experiences. Practitioners express a clear preference for task-oriented, assistive AI systems over end-to-end generative solutions.

AI integrationcreative workflowshuman-AI collaboration

Reverse Engineering of Music Mixing Graphs with Differentiable Processors and Iterative Pruning

Sep 19, 2025
SL
Sungho Lee
🏛️ Seoul National University | Sony AI | Sony Europe B.V.

This work addresses the black-box nature of music mixing by proposing a differentiable graph-structured inverse engineering method to automatically infer the processing chain and combination topology applied to dry source signals from the final mix. The approach constructs a parameterized mixing graph using differentiable audio processors and jointly optimizes both graph structure and parameters via gradient descent, augmented by a dry–wet ratio-guided iterative pruning strategy. Compared to manual modeling, it achieves high-fidelity mix reconstruction (PESQ ≥ 3.2, STOI ≥ 0.92) while removing approximately 67% of redundant processors, substantially reducing model complexity. The framework supports batch-parallel optimization and scalable generation of large mixing graphs. Experiments demonstrate its efficiency, scalability, and perceptual naturalness, establishing a novel paradigm for mixing analysis, AI-assisted production, and interpretable audio processing.

Optimizing differentiable audio processors with gradient descentPruning mixing graphs for efficiency while preserving qualityReverse engineering music mixes to uncover processing techniques

In real-world acoustic scenes, users struggle to manipulate unseparated mixed sound sources. To address this, we propose the first end-to-end, text-driven framework for real-time sound field editing, enabling direct, joint manipulation of multiple concurrent sources—such as “reduce air-conditioner noise and enhance speech”—via natural language instructions, without explicit source separation. Our method integrates large language model–based semantic parsing with a differentiable spectrogram decomposition–filtering–reconstruction architecture, supporting open-vocabulary, zero-shot editing. It is trained on a newly curated dataset comprising 160 hours of audio and 100,000 audio–text pairs. Experiments demonstrate significant improvements in source extraction, suppression, and level control: +3.2 dB in SI-SNR and +0.11 in STOI, while maintaining strong robustness and generalization across complex mixtures with 2–5 overlapping sources.

Control sound sources in mixtures via text instructionsEnhance auditory experience with semantic text filtersRemix multiple sounds simultaneously without separation

Computer Audition: From Task-Specific Machine Learning to Foundation Models

Jul 22, 2024
AT
Andreas Triantafyllopoulos
🏛️ Technical University of Munich | Tampere University | University of Augsburg | Imperial College | Munich Center for Machine Learning | Munich Data Science Institute

To address the longstanding reliance on single-task models and the lack of generalizable representations in computer audition, this paper proposes a systematic framework for constructing Auditory Foundation Models (AFMs). Methodologically, it establishes the core paradigm of AFMs for the first time, integrating unified multi-task modeling, cross-modal (audio–text) aligned representation learning, and instruction-driven human–machine interaction. Technically, the framework encompasses large-scale audio–text contrastive pretraining, multi-task prompt tuning, and self-supervised audio modeling. Experiments demonstrate that the proposed AFM achieves substantial performance gains across 10+ downstream tasks—including automatic speech recognition, sound source separation, and environmental sound classification—while enabling zero-shot transfer and open-domain speech understanding. This work advances computer audition toward generality, multi-task synergy, and natural human–machine interaction.

Consolidate multiple audio tasks into single foundation modelsLeverage cross-modal knowledge for general-purpose audio understandingTransition from task-specific models to auditory foundation models

Latest Papers

What's happening recently
View more

This work addresses the lack of a unified, high-quality transcoding method for spatial audio across diverse acquisition formats—such as Ambisonics or microphone arrays—and arbitrary playback systems. The authors propose a general parametric framework that estimates spatial metadata of primary sources and ambient sound in the time–frequency domain, constructs a spatial covariance model tailored to the target playback setup, and derives an optimal linear downmix matrix. This approach supports independent rotation between acquisition and playback geometries and, for the first time, unifies processing for both Ambisonics and raw microphone array inputs. It accommodates arbitrary array configurations, variable numbers of sources, and arbitrary angular power distributions of ambient sound. Listening tests demonstrate that the method significantly outperforms existing parametric renderers across various content types and playback configurations, with particularly notable perceptual improvements for low-order or geometrically constrained arrays.

Ambisonicsmicrophone arraysspatial audio

Existing audio generation evaluation methods struggle to simultaneously address the industrial sound design requirements of reference guidance, controllable variation, perceptual consistency, and workflow efficiency. This work proposes the first production-oriented evaluation framework for sound effect generation, structured around nine core production criteria and a two-stage protocol that enables systematic, goal-aligned comparison of heterogeneous generation and editing approaches. The framework integrates objective metrics—including Fréchet Audio Distance (FAD), ImageBind-based reference alignment, and diversity scores—with human listening experiments to holistically assess perceptual identity preservation and transient fidelity. Empirical results reveal complementary strengths across baseline methods, with AudioX achieving the best trade-off between reference alignment and output diversity while effectively supporting sound morphing tasks.

audio variationindustrial audio designproduction-oriented evaluation