SAGE: Safety-Aligned Gradient Enforcement for Human--Robot Collaboration

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出SAGE方法,通过安全对齐梯度强制解决多人机器人协作中的决策可解释性和安全性问题,提高任务成功率和减少碰撞。
📝 Abstract
Multi-party human-robot collaboration poses a dual challenge: robot decisions should remain interpretable and auditable, while executed actions must satisfy safety constraints during physical interaction. Combining explainable decision-tree policies with control-barrier-function (CBF) filtering provides a promising architecture but creates two learning mismatches in multi-agent reinforcement learning. Safety projection changes the action applied to the environment, while the coupled proposal graph can misalign independently optimized actor updates with a team-level update. We present safety-aligned gradient enforcement (SAGE) to address both mismatches. Its shield-annealed internalization layer (SAIL) uses a differentiable finite-penalty proposal map while retaining the exact CBF quadratic program for execution, preserving constraint-normal sensitivity to internalize repeatedly active safety constraints. Team-averaged Lyapunov policy optimization (TALO) constructs a team-aware update reference and applies a Lyapunov half-space correction to regulate independent actor updates. Physical experiments with two humanoid robots and a human partner demonstrate deployment feasibility. Across nine simulation scenarios, SAGE achieves a 71.0% success rate with 0.5 collision steps per thousand environment steps. Ablations show that direct CBF filtering reduces collision frequency by 98.5% but decreases success from 67.3% to 59.3%. SAIL reduces proposal violation by 48.8% and proposal-execution correction by 85.2%, while TALO reduces the update-consistency gap by 50.8%.
Problem

Research questions and friction points this paper is trying to address.

multi-party human-robot collaboration
safety constraints
explainable decision-tree policies
control-barrier-function (CBF) filtering
reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Safety-Aligned Gradient Enforcement
control-barrier-function
shield-annealed internalization layer
Team-averaged Lyapunov policy optimization
🔎 Similar Papers
Y
Yisen Li
University of Texas at Arlington
H
Hao Zhang
University of Texas at Arlington and Carnegie Mellon University
R
Ruize Geng
Carnegie Mellon University
Y
Yves Tseng
University of Texas at Arlington
Ding Zhao
Ding Zhao
Carnegie Mellon University
Trustworthy AIAI safetyreinforcement learningautonomous vehiclesrobotics
H. Eric Tseng
H. Eric Tseng
Uni. of Texas at Arlington
Automotive Control