🤖 AI Summary
This study addresses the threat of "safety-compliant task hijacking" in LLM-based robots, wherein actions remain physically safe yet violate user authorization. To this end, it formalizes this attack model and proposes MissionPAIR, an adaptive attack framework, alongside AuthGuard-R, a deterministic authorization layer. By constructing a dual-gate architecture that decouples safety from authorization, the system cryptographically binds robot actions to signed task contexts, accompanied by a formal proof of authorization soundness. Across 240 real-world trials involving multiple models, including Claude Haiku and Qwen2.5, the proposed defense successfully intercepted all 109 unauthorized actions and eleven categories of protocol-level attacks with zero false positives.
📝 Abstract
Large language models are increasingly used as high-level planners for mobile robots, robot manipulators, and autonomous vehicles. Recent studies show that these systems can be influenced through malicious text, speech, visual instructions, retrieved documents, and poisoned sensory context. Most defenses ask whether a proposed action is physically safe. This paper studies a different problem: an action may be physically safe and still violate the mission authorized by the user. An attacker may redirect a delivery robot, replace an approved object, extend a robot's operating region, activate an unnecessary sensor, or delay a mission without creating an immediate physical hazard. We call this attack \emph{safety-compliant mission hijacking}. We propose MissionPAIR, an adaptive attack framework that searches for executable plans that pass a safety gate while violating an authenticated mission. We also propose AuthGuard-R, a deterministic authorization layer that binds every executable action to a signed mission, robot identity, object and region scope, current state, time, and input provenance. AuthGuard-R operates with an independent safety gate, giving a dual-gate architecture. We formalize mission policies over robot traces, define security games, and prove authorization soundness, mission non-escalation, replay resistance, robot binding, provenance separation, threshold-approval security, audit-log tamper evidence, and trace-level composition. We report a preliminary cross-model evaluation with Claude Haiku~4.5 and the open-source Qwen2.5~7B planner. Across 240 live attack trials, the planners followed an injected mission deviation in 109 trials; AuthGuard-R rejected all 109 resulting unauthorized actions. A separate hand-constructed suite of eleven protocol- and policy-level attacks was also blocked completely.