WBAG: A Whole-Body and Attached-Geometry Safety Framework for Vision-Language-Action Manipulation

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the collision risks arising from unmodeled whole-body and grasped-object geometries during Vision-Language-Action (VLA) policy deployment by proposing a safety framework based on Control Barrier Functions (CBFs). The method explicitly models the robot's full-body and attached geometries to construct grasp-conditioned safe sets, transforming dynamic geometric representations into differentiable CBF constraints. By minimizing action modifications, it achieves adaptive obstacle avoidance that responds dynamically to changing grasp states. Evaluated on the SafeLIBERO benchmark, the proposed approach attains a 97.38% scene safety rate and a 59.38% safe success rate, significantly outperforming existing baselines.
📝 Abstract
Vision-language-action (VLA) policies have demonstrated impressive capabilities in generalizable robotic manipulation, but their deployment in the real world remains challenging due to potential collisions involving different parts of the robot, manipulated objects, and the surrounding environment. Existing inference-time VLA safety frameworks typically rely on simplified end-effector-centered representations that do not explicitly model the full articulated robot and attached-object geometry. In this paper, we present WBAG, a safety framework that models the robot's whole-body and grasp-dependent attached geometry. WBAG constructs a grasp-conditioned safe set that adapts the protected geometry as objects are grasped, then converts this evolving geometry into differentiable CBF constraints that minimally modify the VLA's native six-dimensional operational-space action for collision avoidance across robot, scene, and attached geometry. On the SafeLIBERO benchmark, a variant of LIBERO augmented with obstacles for safety evaluation, WBAG achieves the best overall safety and safe task success among the evaluated methods under a scene-level safety evaluator that monitors all eligible non-task objects, reaching 97.38\% aggregate Scene Safety and 59.38\% Safe Success.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
Robotic Manipulation
Collision Avoidance
Safety Framework
Whole-Body Geometry
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action
Whole-Body Geometry
Control Barrier Functions
Robotic Manipulation Safety
Differentiable Constraints
💼 Related Jobs
No related jobs found.
S
Samuel Zhen
Department of Computer Science and Engineering, Texas A&M University, College Station, TX 77843, USA
S
Siwon Jo
GRASP Laboratory, University of Pennsylvania, Philadelphia, PA 19104, USA
Yanze Zhang
Yanze Zhang
University of Illinois at Chicago
RoboticsMulti-Agent SystemsAutonomous DrivingRobot LearningMachine Vision
Wenhao Luo
Wenhao Luo
Assistant Professor, University of Illinois Chicago
RoboticsMulti-Robot SystemsMulti-Agent SystemsMachine LearningControl Theory