Score
Design and produce the auditory elements of a project—creating, recording, editing, synthesizing, and processing sounds such as Foley, sound effects, ambiences, and musical cues. Also specify and implement their spatialization, dynamics, mixing, and technical integration into playback pipelines or interactive engines, and analyze acoustic and perceptual properties to meet timing, intelligibility, and aesthetic requirements.
In real-world acoustic scenes, users struggle to manipulate unseparated mixed sound sources. To address this, we propose the first end-to-end, text-driven framework for real-time sound field editing, enabling direct, joint manipulation of multiple concurrent sources—such as “reduce air-conditioner noise and enhance speech”—via natural language instructions, without explicit source separation. Our method integrates large language model–based semantic parsing with a differentiable spectrogram decomposition–filtering–reconstruction architecture, supporting open-vocabulary, zero-shot editing. It is trained on a newly curated dataset comprising 160 hours of audio and 100,000 audio–text pairs. Experiments demonstrate significant improvements in source extraction, suppression, and level control: +3.2 dB in SI-SNR and +0.11 in STOI, while maintaining strong robustness and generalization across complex mixtures with 2–5 overlapping sources.
This study addresses the labor-intensive and repetitive nature of traditional sound design by systematically evaluating the perceptual credibility of procedurally synthesized audio across different film genres. Leveraging an online procedural audio engine, the authors generated canonical sound effects—such as lasers, impacts, airflow, and rocket propulsion—and conducted the first comprehensive comparison of their realism and suitability in live-action versus animated contexts. Results demonstrate that synthesized sounds achieve high perceptual fidelity in narrative and science-fiction scenes but fall short in cartoonish depictions of everyday actions. Through analysis of critical acoustic features and feedback from professional sound designers, the research identifies key optimization targets and provides empirical grounding and practical guidance for integrating procedural audio into cinematic production pipelines.
To address the weak spatial awareness and lack of natural interaction in conventional desktop-based 3D audio design tools, this paper proposes a six-degree-of-freedom (6DoF) spatial audio design framework tailored for extended reality (XR). Implemented on Apple Vision Pro, the system features an augmented reality (AR) audio design interface integrating 6DoF hand–eye tracking, real-time binaural rendering, and cross-modal feedback to enable intuitive manipulation of virtual sound sources. We introduce two novel design paradigms: “embodiment-aware AR sound design” and “audio-visual modality balance in AR GUIs.” A user study with 27 participants—including domain experts and novices—demonstrates that 6DoF interaction significantly improves spatial localization accuracy (+31.2%) and design intuitiveness. The results further identify high-potential application pathways in education, game development, and accessibility.
This study addresses a significant misalignment between current AI tools and the needs of practitioners in high-level narrative sound design. Employing a mixed-methods approach—combining 76 survey responses with semi-structured interviews of 20 industry professionals—the research systematically examines the current applications, challenges, and expectations surrounding AI in sound design workflows. From the perspective of sound designers, the study identifies five core themes for AI integration: context, workflow, potential, risks, and appropriate use, offering concrete design recommendations for developers. Findings indicate that while existing AI tools are effective for rapid, disposable audio tasks, they lack the narrative depth required for cinematic and immersive experiences. Practitioners express a clear preference for task-oriented, assistive AI systems over end-to-end generative solutions.
To address temporal misalignment, semantic inconsistency, and reliance on costly manual timestamp annotations in video-driven Foley sound synthesis, this paper proposes Video2Sound—a fully end-to-end self-supervised framework. Methodologically, it introduces RMS intensity envelopes as unsupervised temporal event cues for the first time, enabling fine-grained control via RMS discretization and a novel RMS-ControlNet that jointly incorporates textual and audio-semantic prompts. The architecture adopts a two-stage pipeline—Video2RMS followed by RMS2Sound—integrating self-supervised learning, ControlNet-guided diffusion modeling, pretrained text-to-audio (T2A) priors, and explicit RMS feature representation. Extensive evaluation demonstrates state-of-the-art performance in audiovisual alignment, impact timing accuracy, dynamic intensity modeling, timbral fidelity, and controllability over sound details. The code, pretrained models, and interactive demo are publicly released.
This work addresses the limitations of traditional audio effect interfaces, which rely on discrete parameters and hinder holistic perception and fluid exploration of sound transformation spaces. The authors propose a perceptually grounded two-dimensional map interface that organizes audio effects into a continuous soundscape, integrating spatial interaction, DAW-style controls, and embedded machine learning. This system uniquely combines interpretable DAW-style parameter control with embedding-based semantic search, enabling continuous browsing, interpolation, and fine-grained editing of audio effects within a unified interface. Experimental results demonstrate that the approach significantly enhances users’ intuitive understanding of and fluency in navigating sound transformation spaces during music creation, production, and performance.
This study addresses the growing demand in digital applications for high-quality, semantically aligned, and contextually relevant AI-generated sound effects that exhibit both diversity and controllability. It presents a systematic review of sound effect generation models from the past five years, encompassing approaches driven by text, visual, audio, and multimodal inputs. By analyzing 30 peer-reviewed papers sourced from Google Scholar, IEEE Xplore, and ACM Digital Library, this work offers the first comprehensive comparison of how different input modalities influence generation performance. The review highlights advances in audio fidelity, semantic alignment, and temporal coherence, while identifying persistent challenges such as temporal synchronization and perceptual consistency. Furthermore, it underscores a significant gap between current objective evaluation metrics and human perceptual judgments.
Existing audio generation evaluation methods struggle to simultaneously address the industrial sound design requirements of reference guidance, controllable variation, perceptual consistency, and workflow efficiency. This work proposes the first production-oriented evaluation framework for sound effect generation, structured around nine core production criteria and a two-stage protocol that enables systematic, goal-aligned comparison of heterogeneous generation and editing approaches. The framework integrates objective metrics—including Fréchet Audio Distance (FAD), ImageBind-based reference alignment, and diversity scores—with human listening experiments to holistically assess perceptual identity preservation and transient fidelity. Empirical results reveal complementary strengths across baseline methods, with AudioX achieving the best trade-off between reference alignment and output diversity while effectively supporting sound morphing tasks.
Existing Foley sound datasets generally suffer from insufficient quality and coarse annotations, hindering data-driven research in classification, retrieval, and synthesis. To address this gap, this work introduces and publicly releases FoleySet—a large-scale Foley dataset comprising 10,000 audio clips meticulously recorded following professional Foley practices. The dataset features a two-tier manual semantic annotation scheme that precisely aligns synchronous sound effects with on-screen human actions, such as footsteps, clothing rustles, and prop manipulations. FoleySet is the first to offer multi-level annotations, standardized formatting, and a permissive Creative Commons license, thereby filling a critical resource void in the field. It provides strong support for Foley-related audio tasks and advances research toward automated audiovisual content production.
This work addresses the limitations of existing tools and creative bottlenecks faced by composers and sound designers in timbral exploration by proposing an evolutionary generative framework that integrates Quality-Diversity (QD) algorithms with supervised discriminative models. The approach employs multi-band specialized Compositional Pattern Producing Networks (CPPNs) to reduce architectural complexity while preserving performance, and leverages the MAP-Elites algorithm to efficiently search an extended behavioral space spanning multiple duration dimensions, thereby uncovering mechanisms for cross-contextual target switching. The system autonomously generates synthetic sounds exhibiting both diversity and novelty across temporal and contextual dimensions, with its creative potential validated through an online explorer and experimental musical applications.