arXiv · 2610.05129
Representation--Behavior Alignment for Explainable Weakly-Supervised Video Anomaly Detection
Abstract
Multimodal Large Language Models (MLLMs) provide a natural way to make video anomaly detection more explainable. However, their final decisions do not always fully use the discriminative information contained in their hidden states, an issue we refer to as representation--behavior misalignment. We decompose this gap into a capacity component that measures discriminative information never aggregated into the readout position, and a directional component that measures the angular mismatch between the optimal and the native normal--abnormal axis at that position. Across multiple video anomaly detection benchmarks and MLLM backbones the directional component dominates, and residual-stream tracing shows that native-axis separability rises sharply in several mid-to-late attention layers. Because both components are governed by attention rather than MLP updates, we propose Representation--Behavior Alignment (RBA), a parameter-efficient method that adapts those layers using video-level labels alone while updating about 0.012\% of the backbone parameters. Experiments on three benchmarks show that RBA improves native-readout performance and better aligns the model's decision direction with discriminative representations, and it produces anomaly decisions and explanations through a single generative process.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chao Huang, Pengfei Wei, Kaige Li, Chengliang Liu, Wei Wang, Wenqi Ren, Xiaochun Cao. 2026-10-04. Representation--Behavior Alignment for Explainable Weakly-Supervised Video Anomaly Detection. https://arxiv.org/abs/2610.05129
Cite the original work for its findings. Save a collection to share your selection of sources.