arXiv · 2607.27109
MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
Abstract
With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 data sources, covering 6 capability categories and 15 evaluation dimensions. Given a model-generated caption, MMAC checks whether it mentions relevant information in the target dimension and whether the mentioned content is consistent with the reference label. We evaluate representative open-source and proprietary AudioLLMs. Results show clear differences across evaluation dimensions, information coverage, and description reliability. We will release the MMAC benchmark and evaluation code.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Weijie Wu, Junbo Li, Lin Li, Jun Fang, Qingyang Hong. 2026-07-29. MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning. https://arxiv.org/abs/2607.27109
Cite the original work for its findings. Save a collection to share your selection of sources.