arXiv · 2609.23113
Search, Ground, Plan: Functional Sufficiency for Task and Motion Planning under Incomplete Scene Knowledge
Abstract
Foundation models (FMs) have expanded task and motion planning (TAMP) to manipulation problems specified through language and visual observations. However, incomplete scene knowledge leaves a critical gap between understanding what the task requires and knowing whether the physical scene can actually realize it. We introduce GRAB-TAMP, an FM-based TAMP framework that searches for scene entities required for task completion, grounds functional roles to valid physical objects, and plans only after a complete joint assignment establishes functional sufficiency. We represent the task through functional roles, relations, and assignment constraints, and incrementally inspect the scene while requirements remain unresolved, verifying candidate objects through semantic, geometric, and relational checks. We evaluate GRAB-TAMP across 32 scene variants spanning Kitchen, Living Room, and Workshop domains. Across 200 feasible trials, our approach achieves 54.0% end-to-end success with 67.3% plan goal coverage. Compared with three FM-based TAMP frameworks under the same execution setting, GRAB-TAMP improves end-to-end success by 25.7 percentage points over the mean baseline. Implementation and evaluation code: https://github.com/Narendhiranv04/GRAB-TAMP
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Narendhiran Vijayakumar, Nav Singhal, Girish Varma, Antony Thomas. 2026-09-19. Search, Ground, Plan: Functional Sufficiency for Task and Motion Planning under Incomplete Scene Knowledge. https://arxiv.org/abs/2609.23113
Cite the original work for its findings. Save a collection to share your selection of sources.