arXiv · 2609.23131
Selective Commitment for Language-Guided Object Retrieval under Partial Observability
Abstract
Language-guided object retrieval under partial observability requires deciding whether to gather more evidence, interact with the scene, grasp a candidate, or abstain. We present a closed-loop framework that coordinates these decisions for retrieving a target specified in relation to a reference container. The framework maintains a persistent joint belief over target identity, container relation, and presence through tracked-object, unobserved-target, and target-absent hypotheses. View-conditioned categorical VLM observations update this belief; conformal grasp eligibility and robot feasibility govern commitment, while finite-horizon belief-space planning selects information-gathering actions. Across five different scenarios, our proposed method succeeds in 19/25 simulation episodes versus 12/25 for the best-performing task-adapted baseline and is the only evaluated policy to achieve at least one success in each scenario. Ablations show that cross-view memory improves success under partial occlusion, while the full system does not consistently outperform simplified variants. Real-robot trials demonstrate closed-loop re-observation and autonomous recovery from injected grasp failures, while injected viewpoint failures end in false defer. Experimental results demonstrate the feasibility of coordinating evidence gathering and selective grasp commitment within a unified framework for retrieval under partial observability.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wonhee Koh, Sushil Samuel Dinesh, Hansol Ko, Shinkyu Park, Eungjoo Lee. 2026-09-19. Selective Commitment for Language-Guided Object Retrieval under Partial Observability. https://arxiv.org/abs/2609.23131
Cite the original work for its findings. Save a collection to share your selection of sources.