arXiv · 2607.23797
Memory for Attention: Language-Conditioned Re-Perception with a Vision--Language--Motion Map
Abstract
A robot carrying a persistent, behavior-annotated map faces two very different planning questions, and its memory answers only one of them well. The spatial-navigation question - how to walk around a room - we address first, and report a negative: building on Vision-Language-Motion Maps (VLMM), a behavior-aware planner cost cuts a planning-time objective by ~35% over 28 AI2-THOR scenes, but under closed-loop execution the benefit nearly vanishes (~4%) and an on-demand vision-language model (VLM) does as well. The resource-allocation question is different: under a limited perception budget, what should the robot attend to right now to keep its own map fresh? Framing re-perception as this attention decision, we show a persistent map's memory - change-history, or even just recency of last sighting - yields the best of the re-perception schedules we test (held-out), while the memoryless category prior is the weakest of them under the sqrt-law schedule -- though under a Whittle index it leads memory until heterogeneity is real. The gain grows with per-instance heterogeneity as a Cauchy-Schwarz bound predicts, tracking Var(sqrt(lambda)), the variance of root-volatility, and reallocates budget toward the important objects the schedule protects; against a real CLIP movability prior it is +21-26%, of which roughly half survives once that prior's saturated scale is calibrated (+7-13%). The map's full combination earns its keep when the task is language-conditioned: told what to keep track of, VLMM grounds the relevant objects (open-vocabulary) and tracks their change (memory), beating a relevance-weighted recency baseline (+2.9% over 26 queries, at full heterogeneity; the ordering reverses when instance rates track category norms) - and a category prior (+9.2%). The map earns its keep not by telling the robot how to walk around a room, but by telling it what to pay attention to.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Dibyendu Ghosh. 2026-09-12. Memory for Attention: Language-Conditioned Re-Perception with a Vision--Language--Motion Map. https://arxiv.org/abs/2607.23797
Cite the original work for its findings. Save a collection to share your selection of sources.