Score
Design, build, and analyze systems that continuously generate audio output in real time, implementing generator architectures that stream audio while accepting and incorporating new instructions asynchronously. Engineer buffering, background-resolution, and seamless transition (e.g., crossfade) mechanisms to mask backend latency and avoid audible interruptions.
This work addresses real-time audio-to-accompaniment generation—i.e., low-latency, high-fidelity synthesis of coherent instrumental accompaniment (e.g., guitar) synchronized with a singer’s streaming vocal input. Methodologically, we propose the first systematic streaming modeling framework, introducing a “future visibility–output block duration” trade-off mechanism to explicitly quantify the intrinsic tension among latency, coherence, and throughput. We employ a streaming-trained Transformer decoder augmented with a lookahead prediction objective, overcoming the coherence limitations of conventional maximum-likelihood training in real-time settings. Experiments demonstrate feasibility of generating high-quality accompaniment under realistic system latency constraints. Our results reveal the critical role of proactive temporal modeling for real-time music generation and establish a new paradigm for interactive AI music systems.
This work addresses the unnatural inter-sentence silences—up to 9.6 seconds—in conventional real-time game commentary systems, which stem from strictly sequential text generation and speech synthesis. To mitigate this latency bottleneck, the authors propose an end-to-end low-latency commentary architecture that parallelizes large language model–based text generation with speech synthesis and incorporates a multi-candidate sentence pre-caching mechanism to enable immediate voice output at utterance boundaries. The proposed approach reduces the average inter-sentence silence duration to 0.3 seconds and improves temporal alignment with professional commentators by over 40%. A user study involving 120 experienced gamers demonstrates that the system significantly enhances the naturalness of speaking rhythm and overall immersion.
Existing audio generation and editing methods often rely on task-specific architectures, resulting in systems that are complex and difficult to scale. This work proposes AudioWeave, the first unified framework that eliminates the need for task-dedicated modules by leveraging a diffusion Transformer architecture. Through joint conditional modeling, factorized positional encoding, and a multi-stage progressive training strategy, AudioWeave supports both text-to-audio generation and six distinct audio editing tasks within a single model. Experimental results demonstrate that AudioWeave achieves performance comparable to specialized models across all tasks, thereby validating the feasibility and potential of unified audio modeling.
In existing digital audio workstations (DAWs), high-level creative intents—e.g., “warm vocal tone”—are difficult to map efficiently onto low-level parameter adjustments, while AI-based audio generators typically produce single-shot outputs without iterative or reversible human-AI co-creation capabilities. Method: We propose NLP-DAW, a framework centered on large language models (LLMs) that translates natural language instructions—including speech and humming—into executable code for direct control of REAPER DAW. It introduces the Model Context Protocol (MCP) to enable context-aware real-time state querying, fine-grained parameter adjustment, and AI-driven beat generation, augmented by atomic scripts and built-in undo functionality for safe, reversible operations. Contribution/Results: Featuring a voice-first, minimalist interface, NLP-DAW demonstrates stable performance across common music production tasks. User evaluation confirms significant reductions in learning overhead and substantial improvements in perceived control, usability, and creative fluency—establishing a novel paradigm for human-AI collaborative music creation.
This work addresses the challenge of real-time streaming audiovisual character generation by simultaneously ensuring speech–text alignment, cross-segment visual consistency, and low-latency constraints. The authors propose a decoupled architecture comprising an LLM-driven coordinator that produces frame-level aligned audio conditions and employs a progress-aware pointer to maintain text–speech synchronization. A joint audiovisual DiT model performs localized bidirectional denoising within short temporal windows, augmented with a sink-token memory mechanism to suppress visual drift. Efficient deployment is achieved through a two-stage distillation strategy. Evaluated on a single H100 GPU, the method achieves real-time performance and outperforms existing baselines in text fidelity, audiovisual synchronization, visual quality, and streaming stability.
提出了一种前端-后端架构,通过流式ASR转录和文本后端LLM实现全双工语音模型的工具调用,保持低延迟交互并提高工具调用性能。
This study addresses the capability degradation caused by sequentially training streaming and few-step sampling for real-time audio-visual generation. To overcome this, we propose parallel training of causal and few-step adapters on a frozen diffusion Transformer backbone. Exploiting the approximate orthogonality of their update directions, the adapters are combined via direct addition based on model merging principles, thereby avoiding the interference inherent in chained fine-tuning without requiring explicit constraints. By integrating block-autoregressive attention with LoRA, the proposed method achieves continuous real-time generation at approximately 26 fps at 480×832 resolution. The resulting image quality is comparable to that of bidirectional teacher models and significantly surpasses conventional baselines.
本文针对实时互动世界的流式音视频生成问题,提出了StreamAV-Bench基准,通过统一评估框架和32个细粒度维度的测试案例来评估13个代表性系统。
Full-duplex speech models suffer from response timeouts and memory inefficiency due to the tail overhead inherent in audio synthesis. This work proposes a runtime optimization mechanism grounded in the model’s native clock, which achieves on-demand allocation and precise replay through GPU orchestration, static memory analysis, and computation graph recording, thereby eliminating scheduling gaps. Furthermore, streaming state management is dynamically optimized to reduce memory footprint. Experimental results demonstrate that the proposed approach accelerates inference by 2.85× and reduces peak memory consumption by 38.8%, ensuring stable operation within a one-second interaction cycle.
本文提出StepAudio 3 Gen,一种基于残差向量量化(RVQ)标记的离散自回归生成模型,用于解决多种音频类型生成问题,包括文本转语音、声音设计等。