arXiv · 2609.21316
NaViRrator: Robot Navigation from Human-Readable Maps through a Learned Visual Route
Abstract
Human-readable maps provide an intuitive interface for specifying robot destinations, but connecting their schematic geometry to egocentric observations remains challenging. We present NaViRRator, a framework that translates user-specified start and goal locations on such maps into navigation instructions for a pretrained vision-and-language navigation (VLN) policy. Its core method, RouteScribe, separates route inference from verbalization by first generating an explicit route scaffold in map-image coordinates, which a pretrained vision-language model (VLM) converts into a navigation instruction. We construct the scaffold with start--goal line conditional flow matching (SGL-CFM), which deforms a straight start--goal waypoint sequence into a map-conditioned route. During execution, the VLN policy receives only the instruction and egocentric observations, while the map and scaffold remain upstream, allowing executor replacement without retraining the map-to-language modules. Real-world experiments show higher success rates and success weighted by path length (SPL) than direct map-to-instruction generation, A*-based scaffolding, and Gaussian-source conditional flow matching. Qualitative results further show clearer salient turns and better preservation of the intended maneuver sequence, supporting route-grounded language as a modular interface between human-readable maps and pretrained navigation policies.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ayun Lee, Jiseon Kim, Giseop Kim. 2026-09-18. NaViRrator: Robot Navigation from Human-Readable Maps through a Learned Visual Route. https://arxiv.org/abs/2609.21316
Cite the original work for its findings. Save a collection to share your selection of sources.