Mind the Refinement Gap: When Safe High-Level Robot Plans Produce Unsafe Executions
This study addresses the safety verification gap in language-driven robotic systems, where high-level planning and safety monitoring overlook implicit action effects during execution. By auditing discrepancies between RoboGuard’s judgments on abstract plans and graph-refined trajectories, it reveals potential failures in tracking completeness assumptions. Methodologically, this work proposes a graph-based trajectory refinement mechanism serving as both a lightweight mitigation strategy and a diagnostic tool for physical AI safety, integrating semantic graph planning, Linear Temporal Logic (LTL) monitoring, and SPINE natural language instruction generation. Experimental results demonstrate that the predicted deviations were successfully reproduced across all 28 controlled cases, thereby validating the effectiveness of the proposed refinement approach.