From Priors to Perception: Grounding Video-LLMs in Physical Reality
Video Large Language Models (Video-LLMs) excel in general video understanding but often base physical judgments on event expectations rather than observations. We find that they not only rationalize physically impossible events, but also incorrectly report expected outcomes in physically plausible yet counter-intuitive scenarios, despite clear contradictory visual evidence. We provide the first unified account of these failures as Semantic Prior Dominance (SPD): semantic expectations override conflicting visual evidence. To diagnose and overcome these failures, we introduce PriorPair, a high-fidelity paired video benchmark covering both conflicts across eight physical categories. Its category-directed construction pairs prior-aligned events with semantically matched but prior-conflicting counterparts. We further propose Physics-Anchored Reasoner (PhyAR), a lightweight training framework. Its Visually Anchored Reasoning Chain grounds physical attribution and judgment in explicit visual observations, while Paired Data Binding strengthens learning of decisive event differences through matched-pair supervision. Standard LoRA fine-tuning with PhyAR substantially improves physical reasoning and reduces prior-driven judgment bias across different Video-LLM backbones without architectural modifications. These gains extend to an external physical benchmark while largely preserving general video understanding.