How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
Cross-view spatial reasoning remains a weak spot for vision-language models (VLMs): they reason in language and discard the fine-grained geometry the task requires. Thinking with images aims to fix this by generating an intermediate thinking-image, but recent work shows the visual evidence in these traces is largely ignored. We therefore ask how to make visual thinking matter, and what kind of visual thinking works best. We ask these questions for unified multimodal models (UMMs) that natively support interleaved image-text generation. For the how, we propose View Dropout (VDrop), a training-time intervention that hides parts of one input view from the answer span while leaving it visible to the thinking-image tokens. This incentivizes the model to make use of the thinking-image when answering, rather than answering based on the input views only. With the thinking-image now being used in answer prediction, we ask which kind of visual thinking works best. We frame this as a Learnability-Informativeness (L-I) tradeoff and compare three thinking-image variants: top-down, panoramic, and point-matching renderings. Trained on synthetic scenes and evaluated on five real-world out-of-domain benchmarks, panoramic visual thinking with VDrop is the only configuration that is simultaneously informative and learnable, and achieves the best out-of-domain generalization.