Score
Operates and configures digital audio workstation (DAW) software to record, edit, arrange, and mix multitrack audio and MIDI, creating session templates, managing routing and bussing, applying plugins and automation, and preparing stems and final exports. Implements signal‑processing chains, synchronizes external hardware and tempo‑based devices, and manages file organization and session backups for reproducible productions.
In existing digital audio workstations (DAWs), high-level creative intents—e.g., “warm vocal tone”—are difficult to map efficiently onto low-level parameter adjustments, while AI-based audio generators typically produce single-shot outputs without iterative or reversible human-AI co-creation capabilities. Method: We propose NLP-DAW, a framework centered on large language models (LLMs) that translates natural language instructions—including speech and humming—into executable code for direct control of REAPER DAW. It introduces the Model Context Protocol (MCP) to enable context-aware real-time state querying, fine-grained parameter adjustment, and AI-driven beat generation, augmented by atomic scripts and built-in undo functionality for safe, reversible operations. Contribution/Results: Featuring a voice-first, minimalist interface, NLP-DAW demonstrates stable performance across common music production tasks. User evaluation confirms significant reductions in learning overhead and substantial improvements in perceived control, usability, and creative fluency—establishing a novel paradigm for human-AI collaborative music creation.
Music creators face significant challenges in seamlessly integrating deep learning models into digital audio workstations (DAWs). Method: We propose a hosted, asynchronous remote processing framework enabling transparent integration. It comprises a VST/AU-compliant DAW plugin that communicates with Gradio endpoints via a lightweight pyharp Python API, supporting real-time audio/MIDI streaming, MIDI event serialization, and inference of joint audio-MIDI annotation models—all with native in-plugin rendering of interactive UIs and results. Contribution/Results: This work introduces the first DAW-native support for general-purpose MIDI generation and multimodal joint audio-MIDI annotation models. It enables zero-context-switch creative AI workflows and achieves bounded end-to-end latency (<200 ms under typical conditions), while significantly improving plugin stability and cross-platform compatibility (macOS, Windows, Linux). The framework lowers the barrier to adopting AI tools in music production and fosters efficient collaboration between music technology developers and creators.
This work proposes a fully hardware-based single-tone audio synthesis method implemented on an FPGA to meet the demanding requirements of professional audio applications for high-precision, low-latency sinusoidal signals. By leveraging digital signal synthesis techniques and optimized digital logic design, the approach efficiently generates highly stable sine waves at specific frequencies directly in hardware and integrates a digital-to-analog conversion interface for real-world audio output. Entirely eschewing conventional software or hybrid implementations, the solution achieves significantly reduced latency and enhanced timing accuracy through pure hardware execution. Experimental results demonstrate that the resulting synthesizer is well-suited for stringent electronic systems requiring exceptional signal fidelity—such as clock synchronization, communication transmission, and embedded control—offering high precision, minimal resource utilization, and strong real-time performance.
This work proposes the first system that integrates mixed reality (MR) with a shared digital audio workstation (DAW) to overcome the limitations of traditional desktop-based DAWs, which hinder natural performance interaction and low-latency remote collaboration. The system enables geographically distributed users to collaboratively manipulate a single DAW instance in real time while freely moving in physical space. It introduces a novel hands-free interaction interface using physical foot pedals to support collaborative loop recording, establishing a new paradigm for music creation in the emerging “music metaverse.” Built upon an MR platform and a networked synchronization architecture, the system was evaluated through qualitative and speculative design methods with 20 musicians, demonstrating significant improvements in interaction freedom and remote collaborative experience, and offering key design insights for future MR-based music production.
AI-driven audio effect modeling faces a fundamental challenge: existing neural black-box approaches fail to accurately replicate the complex signal routing, parameter coupling, and authentic behavior of professional DSP workflows. Method: This paper introduces the first industrial-grade automated data generation framework for digital audio workstations (DAWs), supporting seamless integration of VST/VST3/LV2/CLAP plugins—including advanced features such as sidechaining and frequency splitting—and enabling efficient configuration via a lightweight metadata interface. The framework employs Docker-based containerization, combined with differentiable signal-flow graph simulation and black-box parameter estimation, to achieve hybrid graph blind identification. Contribution/Results: Experiments demonstrate that, under equivalent computational budgets, our method significantly outperforms state-of-the-art differentiable plugin approaches in both effect graph topology recovery and parameter estimation accuracy. The open-source implementation establishes a new paradigm bridging neural audio modeling and real-world DSP practice.
This study examines the tensions between speed, controllability, and creative autonomy arising from the integration of AI and automation tools in music production. Drawing on ethnographic methods—including in-depth interviews and situated observations—the research systematically investigates how professional recording engineers, mixers, and producers actually use and perceive mainstream AI-powered music technologies. It offers the first practitioner-centered account of the sociotechnical conflicts emerging as AI becomes embedded in real-world workflows, revealing that gains in efficiency often come at the cost of diminished creative control. Building on these insights, the study proposes a user-centered design approach to guide the development of future tools that better balance operational efficiency with the preservation of artistic agency.
This work addresses the inefficiencies of the edit-compile-reload cycle in audio DSP development and the loss of parameter bindings caused by structural changes in DSP code. To resolve these issues, the authors propose a dual-mode CLAP plugin compilation system for the Faust language that integrates static ahead-of-time (AOT) compilation with dynamic interpreted execution. The system introduces an innovative address-based parameter identity matching algorithm and a stable slot allocation mechanism, enabling runtime hot reloading while preserving consistent parameter identities and host automation bindings. As the first officially integrated Faust-to-CLAP compilation pathway, the implementation achieves significant gains in development iteration speed and parameter stability with only approximately 2,400 lines of C++/Python code.
Existing automatic music mixing approaches struggle to simultaneously achieve high-quality output and flexible style control. This work proposes Diff2Mix, the first end-to-end mixing system that integrates diffusion-based generative modeling with a differentiable audio mixer. The method enables global mix style guidance through reference audio while allowing users to explicitly adjust audio effect parameters, thereby offering a highly controllable mixing process. Experimental results demonstrate that Diff2Mix achieves state-of-the-art performance in both objective metrics and subjective listening tests, effectively balancing audio quality with editing flexibility.
This work addresses the error-prone and labor-intensive process of manually rewriting differentiable audio processors for real-time deployment, which often lacks formal verification. To bridge this gap, the authors propose ADAC, a compiler that automatically translates trained differentiable audio models into efficient FAUST code via a framework-agnostic intermediate representation, enabling end-to-end deployment. ADAC supports real-time hot-swapping, incorporates built-in stability verification, and facilitates macro-control parameter design, ensuring that the exported audio plugins match the original model’s impulse response within floating-point precision. Experimental results demonstrate that ADAC successfully converts feedback delay network models into fully functional, stable, and reliable real-time audio plugins, with discrepancies reduced to the level of floating-point noise, thereby establishing a robust pathway from research prototypes to production-grade applications.