arXiv · 2609.34148
Geometric Encoding for Spatial Reasoning in Vision-Language Models
Abstract
Vision-Language Models (VLMs) are far more reliable at recognizing what appears in a video than at reasoning about its spatial and temporal properties, such as metric distances, object dimensions, and consistent object identities across frames. We present Geometric Code, a perception-to-geometry pipeline that computes explicit spatial structure from video and supplies it to VLMs as context to augment reasoning. A perception layer segments and classifies objects and recovers depth, camera pose, and intrinsics from monocular RGB video. A deterministic geometric engine then back-projects, merges, and cleans these outputs into a spatial code, including per-object positions, dimensions, counts, inter-object distances, appearance order, and room geometry. The code is serialized into VLMs' prompts, either alongside the video or replacing it entirely. Specifically, there is no component trained or fine-tuned in our approach. On VSI-Bench, augmenting 2B and 4B open models with the spatial code improves average accuracy by +4.1 points over the frames-only baseline, with the largest gains on numeric estimation tasks such as absolute distance (+24.1 points). The results suggest that explicitly computed geometry, delivered through the language channel, recovers spatial competence that small VLMs cannot extract from pixels alone.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Antonio Jun, Haoshui Yu, Zhengyi Lu, Huirong Fu, Yao Qiang. 2026-09-28. Geometric Encoding for Spatial Reasoning in Vision-Language Models. https://arxiv.org/abs/2609.34148
Cite the original work for its findings. Save a collection to share your selection of sources.