SafeStage: Evaluating Safety Before, During, and After Vision-Language-Conditioned Robot Manipulation

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入SafeStage,一个评估机器人操作安全性的生命周期结构基准,在任务执行前、中、后三个阶段识别和分析安全问题。
📝 Abstract
Vision-language-conditioned robot policies integrate perception, language understanding, and control for general-purpose manipulation. However, existing evaluations often focus on task success, isolated physical constraints, semantic refusal, or realized physical damage, providing limited insight into where safety fails during closed-loop manipulation. We introduce SafeStage, a lifecycle-structured benchmark for evaluating manipulation safety before, during, and after task execution. SafeStage contains 97 purpose-built risk scenarios organized into three stages. Initial-State Hazards captures safety-relevant relations that must be resolved before manipulating the target. Execution-Time Safety evaluates unsafe contacts, trajectories, region entries, and object interactions during execution. Final-State Hazards capture unstable or otherwise unsafe conditions remaining after nominal task completion. The benchmark evaluates realized interactions using event-based and state-based checks and reports native task success independently from stage-specific safety outcomes. We evaluate representative direct-action Vision-Language-Action (VLA) policies and policies with world-model-based policies under a common closed-loop protocol. Our results demonstrate that nominal task completion frequently coexists with safety violations and that different policies exhibit distinct failure profiles across the three stages. By separating task success from safety and localizing when violations occur, SafeStage provides a unified diagnostic testbed for evaluating and improving vision-language-conditioned robot manipulation policies.
Problem

Research questions and friction points this paper is trying to address.

safety evaluation
vision-language-conditioned robot manipulation
closed-loop manipulation
risk scenarios
task success
Innovation

Methods, ideas, or system contributions that make the work stand out.

SafeStage
Lifecycle-Structured Benchmark
Vision-Language-Conditioned Robot Manipulation
Safety Evaluation
💼 Related Jobs
No related jobs found.
J
Jinzhu Luo
Worcester Polytechnic Institute, Worcester, MA, USA
Q
Qi Zhang
Worcester Polytechnic Institute, Worcester, MA, USA
W
Wei Wang
Futurewei Technologies Inc., Santa Clara, CA, USA
W
Wei Jiang
Futurewei Technologies Inc., Santa Clara, CA, USA