SafeLoop: Risk-Aware Rollback for Vision-Language-Action Manipulation

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决视觉-语言-动作模型在长时间操作中的脆弱性问题,通过SafeLoop机制预测风险并基于回滚恢复,从而减少约70%的风险案例。
📝 Abstract
Recent vision-language-action (VLA) models are promising for general-purpose manipulation, but long-horizon execution remains fragile. Small state-estimation or control errors can lead to irreversible failures (e.g., collisions and object drops). Avoiding these risks requires a proactive safety mechanism capable of anticipating hazards. In this paper, we introduce SafeLoop, a non-invasive external wrapper that adds hazard prediction and rollback-based recovery to a VLA model without changing its parameters. SafeLoop trains a risk predictor from vision and proprioception to output four values: the probability and time-to-hazard for body collisions and for object failures. A lightweight controller then chooses one of three actions based on the predicted risk: continue execution (noop), save a safety checkpoint (record), or retreat in joint space (rollback). Rollback moves the robot back to a recent safe waypoint and queries the base policy again, which may yield an alternative continuation. Across 24 LIBERO tasks (16 random seeds each) and three real-robot tasks (25 rollouts each), SafeLoop achieves a stronger overall safety-success trade-off than alternative methods, reducing hazard cases by roughly 70% while preserving task success and the base-policy control rate. Project code is available at https://github.com/Loule0-0/SafeLoop/tree/release/safeloop.
Problem

Research questions and friction points this paper is trying to address.

vision-language-action
hazard prediction
long-horizon execution
irreversible failures
safety mechanism
Innovation

Methods, ideas, or system contributions that make the work stand out.

risk prediction
rollback-based recovery
non-invasive external wrapper
safety checkpoint
💼 Related Jobs
No related jobs found.
Z
Zeyu Lou
Nanjing University, Nanjing, China
T
Tianran Zhang
The Hong Kong University of Science and Technology (Guangzhou), Guangdong, China
X
Xinquan Yue
Nanjing University, Nanjing, China
Ya Jing
Ya Jing
ByteDance Research
Computer VisionRoboticsCross-modal Learning
C
Chenyang Si
Nanjing University, Nanjing, China