arXiv · 2609.20849
Enhancing Audio Reasoning via Semantic Summary Prediction
Abstract
Large Audio Language Models (LALMs) perform well on complex question answering but often show a reasoning gap, where explicit Chain-of-Thought (CoT) reduces accuracy compared to direct answers. We hypothesize that long reasoning sequences shift attention away from the audio input. To address this, we propose SPARE (Semantic Prediction for Audio REasoning), which introduces a register token aligned with the final conclusion using a cosine similarity loss with a Sentence-BERT embedding. This conditions the model's latent space with the target semantic goal before reasoning begins. Experiments on MMAU and MMAR with SALMONN show improved zero-shot reasoning and stronger early attention to audio without additional inference cost.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Francesco Bonzi, Pooneh Mousavi, Cem Subakan, Mirco Ravanelli. 2026-08-07. Enhancing Audio Reasoning via Semantic Summary Prediction. https://arxiv.org/abs/2609.20849
Cite the original work for its findings. Save a collection to share your selection of sources.