arXiv Science⌕ Search

arXiv · 2609.22260

Causal Localization of the Refusal Direction in Audio Language Models

Abstract

A large audio language model (LALM) attaches a speech front end to a text language model (LM) that is already safety-aligned. When such a model refuses a harmful spoken request, is the refusal carried by the front end, or inherited from the text LM? We test this with causal interventions. At each model's audio-to-LM interface and at tested LM residual layers, we fit a direction separating harmful from benign prompts, ablate its component, and measure the resulting change in the model's first-token refusal margin. Four of the five models are evaluated under held-out category shift. Across five LALMs spanning three backbone families, with three models passing a baseline safety gate, the largest tested effects occur in a mid-to-late LM band, while ablations of the tested interface directions have little effect. On Qwen2.5-Omni, ablating the L16 direction changes the margin by -7.10, versus -0.013 at the projector. The audio pathway is still in use: zeroing the encoder output changes the margin by -4.7. In the same model, the contrast is linearly decodable at an early layer where single-layer ablation has little effect, and a direction fitted on the LM backbone alone transfers to the full audio model. Because harmful and benign prompts also differ in form, we interpret the direction as refusal-linked rather than harmfulness-specific. These interventions localize dependence of the refusal margin, not where refusal is computed. Moving the margin also does not always change what the model writes. Safety audits of these models should use interventions rather than rely on probes alone, and should examine the inherited text LM alongside the audio interface.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Leonardo Haw-Yang Foo, Hung-yi Lee. 2026-09-06. Causal Localization of the Refusal Direction in Audio Language Models. https://arxiv.org/abs/2609.22260

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

FATE: Frame-Level Audio-Visual Temporal Embedding

When a dog opens its mouth and barks, humans naturally recognize what the sound is and when it occurs. Building audio-visual models with this same ability requires representations that capture both semantic and temporal alignment. Current approaches fall short on one side or the other: embedding models match semantic but lose temporal information; synchronization models capture temporal offsets but lack semantic understanding. To bridge this gap, we propose FATE, Frame-level Audio-visual Temporal Embedding. Unlike prior embedding models that pool each modality into a single embedding and discard temporal information, FATE retains frame-level sequences, aligns them on the physical timeline, and computes similarity over strictly aligned frame pairs. Unlike synchronization models that output only an offset prediction, FATE encodes synchronization in a reusable embedding space, trained with a joint objective combining cross-video semantic and within-video temporal contrastive learning to capture both what sounds and when it occurs. Across three tasks, FATE surpasses the strongest baseline on temporal and semantic retrieval by a large margin, matches fully supervised methods on event localization in a zero-shot setting, and achieves the best correlation with human judgments as a generation evaluation metric. The source code can be found at \texttt{https://github.com/guankaisi/FATE}.

cs.MM↗

Emotion Understanding in Streaming Video with Trajectory-Aware Reliability

Video emotion understanding is commonly studied as an offline classification problem, where the complete video segment is available before prediction. Real-time interaction, however, requires emotion decisions from incomplete and evolving evidence. This paper studies streaming video emotion understanding as a reliability-aware decision process over evolving emotion beliefs. In this setting, a single confident prefix prediction can still be unreliable when the underlying belief trajectory is unstable or repeatedly switches across emotion classes. We propose TRACE, a trajectory-aware reliability framework that forms low-latency emotion beliefs from streaming audio prefixes, estimates reliability from confidence, entropy, stability, and class-switching patterns, and selectively invokes contextual belief reinterpretation with visual, textual, and neighboring-utterance evidence. TRACE keeps stable cases in the low-latency online pathway while allocating stronger multimodal reasoning to uncertain cases that remain ambiguous. Experiments on StreamMER, MELD, and MER2024 show that TRACE improves the accuracy-cost trade-off, retaining most full-context gains while reducing unnecessary contextual reasoning.

cs.MM↗

From Expression to Reaction: Role-aware Visual Transfer and Stimulus-guided Reasoning for Interlocutor Emotion Recognition

In this paper, we propose a Role-aware Stimulus-guided (RASG) framework for interlocutor emotion recognition, which predicts listener emotions from listener-only videos and speaker-only audios. RASG consists of Role-aware Visual Transfer (RVT) and Stimulus-guided Boundary Reasoning (SBR) modules, which address supervision mismatch due to the lack of labeled listener data and ambiguity among visually similar listener reactions whose interpretation depends on speaker context, respectively. More specifically, RVT selects speaker samples whose facial expressions support their emotion labels. It then filters listener tracks and uses reliable pseudo-labels to train a listener-centric visual expert. SBR uses a two-class language reasoner only when the visual model is uncertain. It treats speaker audio and text as context rather than direct emotion evidence to distinguish similar listener reactions. Experiments conducted on MER-Cross dataset shows that RASG achieves 76.25\% on MER-Cross and improves the performance of the baseline over 17\%. Our team ranks second in Track 1 (MER-Cross) of the MER Grand Challenge at ACM MM 2026.

cs.MM↗