arXiv · 2608.28212
MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding across Score and Performance
Abstract
Musicians commonly communicate music through scores and performances. Scores encode musical intent, while performances realize it in sound. To investigate whether models can meaningfully engage with both modalities, we introduce MuSP-Bench, a human-authored benchmark of 490 questions targeting understanding across Musical Scores and Performances. The benchmark distinguishes itself by spanning score-based, performance-based, interpretive, and long-horizon reasoning across classical piano and orchestral works. We evaluate frontier multimodal large language models under multiple input conditions. Our results show that these models struggle substantially to understand scores, while facing even greater challenges when reasoning about performance audio. The benchmark is available at https://musp.vaclis.net/.
Explore related subjects
Keep this discovery
Milan Liessens Dujardin, Song-Ze Yu, Kevin Miao. 2026-08-28. MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding across Score and Performance. https://arxiv.org/abs/2608.28212
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.