arXiv · 2609.26028
REVE: Efficient Hallucination Correction for Large Audio-Language Models via Reused Encoder States
Abstract
Large audio-language models may mention acoustic events that are absent from the input. A separate audio event detector can verify these mentions, but doing so requires a second audio encoder and a separate forward pass. We propose Reused Encoder States for Verifying Events (REVE), a lightweight method that uses states already computed by the target model. One readout summarizes class scores across audio frames, while another uses pooled states from four consecutive frame intervals. Class-aware score fusion combines their outputs to verify generated event mentions without encoding the audio again. On AudioSet, REVE removes 92.9% of label-unsupported mentions under a faithful-mention recall constraint. With fewer added parameters and no second audio-encoding pass, REVE achieves a reduction comparable to those of CED-Tiny and CED-Base. Its complete verification latency is about 1/18 of the CED-Base path. Results on controlled DESED mixtures and different target-model architectures further confirm the effectiveness of encoder-state reuse.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hongjin Song, Jiasheng Kuang, Xinyu Yang, Qiuyu Fang, Ziyu Wu, Guowu Tan, Xiang Xie. 2026-09-22. REVE: Efficient Hallucination Correction for Large Audio-Language Models via Reused Encoder States. https://arxiv.org/abs/2609.26028
Cite the original work for its findings. Save a collection to share your selection of sources.