PlanGuard: A Guardrail for Multi-Step Plan Safety in Embodied Agents

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the overlooked problem of compositional physical risks in multi-step planning for embodied agents by proposing PlanGuard, the first pre-execution physical safety verification framework. The method introduces a novel STAC-OPD distillation technique that integrates supervised fine-tuning (SFT), on-policy distillation, token-level distribution transfer, and probabilistic routing sequence compensation to enable compact models to adaptively compensate for strong teacher supervision. Experimental results demonstrate that PlanGuard, with only 2B parameters, achieves an average accuracy of 87.15% and an F1 score of 87.21% on the test set, effectively realizing comprehensive safety assessment of complete multi-step plans within environments.
📝 Abstract
Embodied task planners may produce multi-step plans whose subtask dependencies and interactions with the environment create physical risks during execution. Yet existing safeguards overlook such compositional risks, as general-purpose guardrails focus on semantic harm and embodied safety detectors assess subtasks in isolation. To address this gap, we introduce PlanGuard, the first pre-execution detector that evaluates the physical safety of a complete multi-step plan in its current environment. For training and evaluation, we construct a Multi-Step Plan Safety (MSP-Safe) dataset through paired task construction, plan generation using diverse planners, and safety annotation by three judges. Task-oriented SFT on MSP-Safe establishes fundamental plan-safety assessment capabilities, yet a substantial gap remains between compact models suitable for real-time deployment and stronger but costlier large models. Accordingly, we propose Strong-Teacher Adaptive Compensation for On-Policy Distillation (STAC-OPD), which provides compact models with adaptive strong-teacher supervision along their on-policy trajectories. It combines token-level distribution transfer from a fine-tuned strong teacher with probability-routed sequence-level compensation, retaining student-generated targets when the student favors the reference safety decision and using teacher-reconstructed targets otherwise. Across all test subsets, PlanGuard-2B achieves average 87.15% ACC and 87.21% F1, demonstrating effective whole-plan physical-risk detection at compact model scale. Code and dataset will be publicly released.
Problem

Research questions and friction points this paper is trying to address.

embodied agents
multi-step plan safety
physical risk detection
guardrail
compositional risk
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Step Plan Safety
Embodied Agents
Knowledge Distillation
On-Policy Distillation
Safety Guardrail
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Junchi Chen
Ant Digital Technologies, Ant Group
Changtao Miao
Changtao Miao
University of Science and Technology of China
AI
Y
Yuxiao Xiang
Anhui Province Key Laboratory of Digital Security
Zhenchao Jin
Zhenchao Jin
USTC > HKU
computer visionmachine learninginformation security
Haojie Yuan
Haojie Yuan
Individual Researcher
Qi Chu
Qi Chu
University of Science and Technology of China
Computer visionArtificial intelligence security
T
Tao Gong
Anhui Province Key Laboratory of Digital Security
H
He Liu
Ant Digital Technologies, Ant Group
B
Bo Zhang
Ant Digital Technologies, Ant Group
J
Jiansheng Cai
Ant Digital Technologies, Ant Group
Z
Zhe Li
Ant Digital Technologies, Ant Group
Nenghai Yu
Nenghai Yu
University of Science and Technology of China
Computer VisionArtificial IntelligenceInformation Hiding