Multi-Link Safety Filtering for VLA Policies Around Moving Hazards

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of vision-language-action policies to collisions with non-target objects in cluttered environments, which hinders their safe real-time deployment. To overcome this, we propose a training-free multi-link safety filter that introduces the first protective mechanism encompassing the entire manipulator arm rather than solely its end-effector. The approach leverages a five-ellipsoid geometric representation and control barrier functions, integrated with RGB-D perception and sparse optical flow, to enable efficient dynamic obstacle tracking and filtering. Experimental results demonstrate that the proposed method reduces the collision rate to 27.27% and improves the safe success rate to 50.43% in simulation. Real-world evaluations further confirm significantly fewer physical contacts with an additional CPU latency of merely 2.2 ms, achieving effective safety assurance under lightweight edge computing constraints.
📝 Abstract
A vision-language-action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show that the policy is safe to deploy in clutter. We study how to keep a pretrained VLA policy clear of such hazards at run time without retraining it, which requires guarding more of the arm than the end effector, following the hazard as it moves, and sharing onboard compute with the policy. Our training-free shield covers the gripper, wrist, and forearm with five ellipsoids and filters every commanded motion through one barrier program against a keep-out ellipsoid fitted from RGB-D perception at reset. Sparse optical flow then carries that ellipsoid's center along with the hazard, with no repeated detection or refitting. Over six simulated hazard-motion conditions, the shield lowers collision from $65.62\%$ to $27.27\%$ and raises safe-success, task completion without collision, from $29.35\%$ to $50.43\%$. Ablations show that guarding the arm links protects beyond end-effector shielding, and that tracking recovers most of the protection lost when the hazard estimate is frozen at reset. On heterogeneous edge hardware, the five-ellipsoid barrier runs on the CPU in $2.2$~ms at the 99th percentile, and trimming the vision--language prefix and taking fewer flow-matching steps shortens each $π_{0.5}$ policy call on the integrated GPU from $343$ to $177.3$~ms. On a physical SO-101 arm across four tasks, the arm touched the hazard in 3 of 16 shielded episodes versus 11 of 16 unshielded ones. Project page: https://yathag.github.io/multilink-safety-filter/
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action policy
Safety filtering
Moving hazards
Collision avoidance
Multi-link protection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Policy
Safety Filtering
Control Barrier Function
Sparse Optical Flow
Training-free Shield
🔎 Similar Papers
No similar papers found.