AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing text-to-audio generation evaluation metrics, which rely on global similarity measures and fail to capture fine-grained semantic errors. To this end, the authors propose the first structured, complexity-aware evaluation benchmark that characterizes soundscape complexity through dimensions such as modality-aware semantic structure, event density, and structural complexity. The framework introduces a fine-grained binary question-answering scoring mechanism grounded in audio anchoring, enabling scalable and interpretable quality analysis. Experiments across 13 open-source models reveal significant deficiencies in current methods regarding attribute control, speech preservation, and compositional soundscape generation. Human evaluations confirm that the proposed benchmark aligns more closely with human semantic judgments than conventional metrics.
📝 Abstract
Text-to-audio (TTA) generation has recently achieved remarkable progress in synthesizing realistic audio from natural language descriptions. However, determining whether generated audio faithfully satisfies complex textual instructions remains challenging. Existing benchmarks mainly rely on global similarity metrics, providing limited insight into fine-grained semantic failures. To address this limitation, we introduce \textbf{AudioScape-TTA}, a structured and complexity-aware benchmark for fine-grained TTA evaluation. AudioScape-TTA represents realistic soundscapes through modality-aware semantic structures and characterizes generation complexity using event density and structural complexity. Based on these annotations, we propose a rubric-based audio-grounded evaluation framework that verifies event realization, acoustic attributes, and speech content through fine-grained semantic criteria. The benchmark contains 2,258 audio-text pairs with 25,707 binary QA rubrics, enabling scalable and interpretable analysis of TTA systems. Experiments on 13 representative open-source TTA models reveal persistent limitations in fine-grained attribute control, speech-content preservation, and compositional soundscape generation. Human validation further demonstrates that our rubric-based evaluation achieves stronger alignment with human semantic judgments than conventional global similarity metrics.
Problem

Research questions and friction points this paper is trying to address.

text-to-audio
fine-grained evaluation
semantic fidelity
soundscape generation
evaluation benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

structured soundscape
fine-grained evaluation
rubric-based assessment
text-to-audio generation
semantic alignment
🔎 Similar Papers
No similar papers found.