arXiv · 2610.10208
CARES: A Controlled Synthetic Benchmark of Speaker Reactions to Sound
Abstract
Automatic audio scene description turns a recording into a text account of a situation. One difficulty is deciding which elements of the audio should be kept, since a description cannot include them all. Annotators disagree about this, making a ground truth hard to obtain. In this work, we first define the ground truth, then generate the data. We focus on audio events and define sound salience with a simple rule: a sound is salient when a speaker audibly reacts to it. For scale and variety, a controlled set of scenarios fixes the ground truth, and a language model writes the dialogues. The resulting corpus, CARES, contains 10,000 two-speaker scenes. We then benchmark six audio-language models on three tasks: identifying the scene, tagging the sounds present, and classifying reactions. We show that these models hear the sounds but miss how the speakers react to them.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Marcel Gibier, Thomas Thebaud, Olivier Boëffard, Jean-François Bonastre. 2026-10-07. CARES: A Controlled Synthetic Benchmark of Speaker Reactions to Sound. https://arxiv.org/abs/2610.10208
Cite the original work for its findings. Save a collection to share your selection of sources.