arXiv · 2610.02876
Seeing, Saying, but Not Using: From Reportable Spatial Facts to Usable States in Multimodal Large Language Models
Abstract
A multimodal large language model that correctly reports a spatial fact does not necessarily use that fact in subsequent reasoning. To study this distinction, we introduce \textsc{SpaceConflict}, a benchmark of 23{,}196 inputs for the construction and use of spatial state. Under a unified Supported/Contradictory/Unknown judgment interface, it covers local fact binding (L1), relational composition (L2), cross-observation consistency (L3), and state judgment under transformation (L4). Posing a direct-state query, a full-transformation query, and an explicit-initial-state query on the same world reveals an availability--utilization gap: models recover the initial state from visual evidence yet fail when that state must drive a transformation. For Qwen3.5-9B, 50 of 100 sequences with a correctly recovered initial state fail the full transformation, and supplying the state explicitly repairs all 50; the gap narrows with scale but does not close. We therefore propose Operational State Supervision (OSS), which supervises task-relevant spatial states and their transformation trajectories and aligns shared facts across contexts. OSS improves paired accuracy on matched judgments most on L3 and L4, the levels that depend on organizing and using state. Evaluating multimodal spatial reasoning thus requires asking not only whether a model can see and state a spatial fact, but whether that fact becomes a usable state in subsequent computation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jinchang Zhang, Guoyu Lu. 2026-10-02. Seeing, Saying, but Not Using: From Reportable Spatial Facts to Usable States in Multimodal Large Language Models. https://arxiv.org/abs/2610.02876
Cite the original work for its findings. Save a collection to share your selection of sources.