🤖 AI Summary
Existing audio generation models lack comprehensive evaluation across multiple dimensions—such as semantic fidelity, speaker consistency, and temporal control—in the context of mixed audio. To address this gap, this work introduces the first fine-grained, multi-control benchmark for mixed audio generation, comprising approximately 4,000 human-verified speech, music, and sound effect segments with precise temporal alignments, along with dedicated subsets for voice cloning and temporally conditioned generation. Leveraging multimodal human annotations and a multidimensional evaluation protocol—including acoustic fidelity, speech quality, semantic alignment, and temporal accuracy—we systematically assess state-of-the-art models. Our evaluation reveals significant trade-offs among these capabilities, with no single model achieving consistent superiority across all dimensions, thereby highlighting the core challenges in controllable mixed audio generation.
📝 Abstract
Recent audio generation systems have progressed from single-modality synthesis to generating complex acoustic scenes containing speech, music, and sound effects. Therefore, evaluating these models requires assessing multiple interacting capabilities, including semantic fidelity, speaker consistency, and temporal control, yet existing benchmarks focus on isolated domains or coarse-grained descriptions. To address this gap, we introduce the Multi-control Mixed Audio Generation (MMAG) benchmark. MMAG contains approximately 4,000 manually verified audio clips with rich annotations covering speech content, speaker identity, music attributes, sound events, and temporal relationships, together with dedicated subsets for voice cloning and timestamp-conditioned generation. We further propose a systematic evaluation protocol that measures acoustic fidelity, speech quality, semantic alignment, and temporal accuracy. Benchmarking representative agentic orchestrators, unified audio-visual generation models, and native mixed-audio generators reveals substantial performance trade-offs across these capabilities, with no existing model performing consistently well. Our results highlight the remaining challenges of controllable mixed audio generation and establish MMAG as a comprehensive benchmark for future research.