arXiv · 2609.39150
OP-CAD: On-Policy Clean-Audio Distillation for Robust Audio-Visual Reasoning
Abstract
Omni-modal large language models deployed in real-world environments encounter external noise that can interfere with their perception and understanding of multimodal inputs. We study their robustness in audio-visual understanding, focusing on question answering under environmental noise and competing speech. The challenge is to resist acoustic interference while preserving useful audio evidence. On-policy distillation provides dense teacher feedback on student-generated responses, but uniform token weighting does not explicitly prioritize positions affected by acoustic interference. We introduce OP-CAD (On-Policy Clean-Audio Distillation), a curriculum-based privileged self-distillation framework for robust audio-visual understanding. Training progresses from mild to severe environmental noise and competing speech, with selective token-level supervision at each stage. The student generates responses from corrupted audio-visual input, while a frozen teacher uses clean audio and the verified answer to supervise the same response prefixes. To allocate this supervision, OP-CAD compares teacher predictions under clean, corrupted, and visual-only contexts without revealing the answer. These matched comparisons measure sensitivity to audio removal and corruption; a bounded weighting rule emphasizes positions identified by either signal while retaining supervision throughout the response. OP-CAD outperforms the compared methods across all evaluated noise conditions. Paired analyses further show improved preservation of clean-correct answers under strong interference, with no observed aggregate clean-accuracy penalty. These results demonstrate the value of directing clean-teacher supervision toward acoustically sensitive predictions for robust audio-visual reasoning.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xingming Shui, Dapeng Chen, Bowei Liu, Jingqi Tian, Minfu Li, Kun Yi, Jiapeng Hong, Yansong Tang. 2026-09-30. OP-CAD: On-Policy Clean-Audio Distillation for Robust Audio-Visual Reasoning. https://arxiv.org/abs/2609.39150
Cite the original work for its findings. Save a collection to share your selection of sources.