π€ AI Summary
This study addresses the safety verification gap in language-driven robotic systems, where high-level planning and safety monitoring overlook implicit action effects during execution. By auditing discrepancies between RoboGuardβs judgments on abstract plans and graph-refined trajectories, it reveals potential failures in tracking completeness assumptions. Methodologically, this work proposes a graph-based trajectory refinement mechanism serving as both a lightweight mitigation strategy and a diagnostic tool for physical AI safety, integrating semantic graph planning, Linear Temporal Logic (LTL) monitoring, and SPINE natural language instruction generation. Experimental results demonstrate that the predicted deviations were successfully reproduced across all 28 controlled cases, thereby validating the effectiveness of the proposed refinement approach.
π Abstract
Language-enabled robot systems increasingly combine semantic-graph planning with temporal-logic safety monitors. We investigate a trace-completeness assumption in these systems: whether the high-level action sequence checked by a monitor represents the navigation and implicit action effects induced during execution. We audit this assumption in RoboGuard by comparing its verdict on a surface plan with its verdict on a graph-refined trace under the same Linear Temporal Logic (LTL) specification. Our evaluation comprises 28 controlled cases spanning five action-abstraction families and 14 end-to-end cases in which SPINE [1] generates plans from natural-language instructions while RoboGuard generates scene-grounded safety specifications. In the controlled evaluation, all 12 targeted abstraction cases exhibit the predicted surface-versus-refined discrepancy while all 16 controls behave as expected, motivating graph-based trace refinement as a lightweight mitigation and a diagnostic tool for physical-AI safety monitors.