arXiv · 2610.02662
Mind the Refinement Gap: When Safe High-Level Robot Plans Produce Unsafe Executions
Abstract
Language-enabled robot systems increasingly combine semantic-graph planning with temporal-logic safety monitors. We investigate a trace-completeness assumption in these systems: whether the high-level action sequence checked by a monitor represents the navigation and implicit action effects induced during execution. We audit this assumption in RoboGuard by comparing its verdict on a surface plan with its verdict on a graph-refined trace under the same Linear Temporal Logic (LTL) specification. Our evaluation comprises 28 controlled cases spanning five action-abstraction families and 14 end-to-end cases in which SPINE [1] generates plans from natural-language instructions while RoboGuard generates scene-grounded safety specifications. In the controlled evaluation, all 12 targeted abstraction cases exhibit the predicted surface-versus-refined discrepancy while all 16 controls behave as expected, motivating graph-based trace refinement as a lightweight mitigation and a diagnostic tool for physical-AI safety monitors.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Stabak Das, Priyesh Ranjan, Xiangfang Li, Lijun Qian. 2026-10-02. Mind the Refinement Gap: When Safe High-Level Robot Plans Produce Unsafe Executions. https://arxiv.org/abs/2610.02662
Cite the original work for its findings. Save a collection to share your selection of sources.