arXiv · 2603.05128
PolyBench: A Benchmark for Compositional Reasoning in Polyphonic Audio
Abstract
Large Audio Language Models (LALMs) are increasingly capable of reasoning over audio, yet existing benchmarks offer limited coverage of reasoning in polyphonic audio, where multiple sound events co-occur and induce compositional structure. To address this gap, we introduce PolyBench, a benchmark designed to evaluate compositional reasoning in polyphonic audio, comprising five evaluation subsets that cover counting, classification, detection, concurrency, and duration estimation, all of which require reasoning over multiple concurrent events and their relations. Our evaluation of state-of-the-art LALMs reveals consistent performance degradation in polyphonic settings, indicating a fundamental bottleneck in current LALMs.
Explore related subjects
Keep this discovery
Yuanjian Chen, Yang Xiao, Han Yin, Xubo Liu, Jinjie Huang, Ting Dang. 2026-03-05. PolyBench: A Benchmark for Compositional Reasoning in Polyphonic Audio. https://arxiv.org/abs/2603.05128
Cite the original work for its findings. Save a collection to share your selection of sources.