MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of fine-grained diagnostic capability in existing audio captioning evaluation methods, particularly regarding information coverage and description reliability. To bridge this gap, we propose the first unified evaluation framework specifically designed for audio captioning that jointly assesses both coverage and faithfulness. We further introduce MMAC, a large-scale, multidimensional benchmark encompassing six core capability categories and fifteen fine-grained evaluation dimensions. MMAC leverages multi-source audio integration, a dimension-aware annotation schema, consistency verification mechanisms, and an automated evaluation pipeline to precisely measure the relevance and accuracy of model-generated captions across all dimensions. Extensive experiments on multiple open- and closed-source audio language models validate the effectiveness of our benchmark, revealing nuanced performance differences. The dataset and code will be publicly released.
📝 Abstract
With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 data sources, covering 6 capability categories and 15 evaluation dimensions. Given a model-generated caption, MMAC checks whether it mentions relevant information in the target dimension and whether the mentioned content is consistent with the reference label. We evaluate representative open-source and proprietary AudioLLMs. Results show clear differences across evaluation dimensions, information coverage, and description reliability. We will release the MMAC benchmark and evaluation code.
Problem

Research questions and friction points this paper is trying to address.

audio captioning
evaluation benchmark
information coverage
description reliability
multi-dimensional assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

audio captioning
multidimensional benchmark
AudioLLMs
information coverage
description reliability
🔎 Similar Papers
No similar papers found.