Adversarial Stress Testing of SPARK Humanoid Safety Filters

📅 2026-05-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the critical challenge of safely deploying humanoid robots, whose high-dimensional dynamics and dense interactions render them particularly vulnerable to failures under complex disturbances. While existing safety filters lack systematic robustness evaluation in such settings, this study proposes an adversarial stress-testing framework that reproduces the SPARK benchmark scenarios in MuJoCo and integrates multiple safety methods—including RSSA, RSSS, SSA, CBF, PFM, and SMA—under challenging conditions such as obstacle-dense environments, distance noise, and information delay. A post-processing pipeline extracts multidimensional metrics, including tracking error, minimum distance, and collision steps, revealing failure modes obscured by conventional benchmarks. Experiments demonstrate inherent trade-offs between tracking accuracy and collision avoidance across methods, with significant degradation in safety performance under perturbations, thereby underscoring the necessity of rigorous robustness validation prior to real-world deployment.
📝 Abstract
Humanoid robots are difficult to deploy safely because they have high-dimensional bodies, many collision constraints, and must operate near people and obstacles. Safety filters help by modifying a nominal control action when it may violate collision-avoidance constraints. Still, nominal benchmark scores do not fully show how these filters behave in harder environments. In this work, we study the robustness of SPARK humanoid safety filters through replication and stress testing. We replicate the SPARK benchmark case G1SportMode_D1_WG_SO_v1 in MuJoCo and evaluate RSSA, RSSS, SSA, CBF, PFM, and SMA under controlled random seeds. We also built a post-processing pipeline that converts raw SPARK logs into goal-tracking, minimum-distance, and collision-step metrics. Our results show that some methods track the goal more closely, while others reduce collision steps more effectively. The stress tests further indicate that safety behavior can change under obstacle crowding, noisy distance estimates, and delayed obstacle information. These findings suggest that humanoid autonomy should be evaluated beyond nominal performance, using metrics that expose failure modes before deployment.
Problem

Research questions and friction points this paper is trying to address.

humanoid robots
safety filters
adversarial stress testing
collision avoidance
robustness evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

adversarial stress testing
humanoid safety filters
robustness evaluation
collision avoidance
failure mode analysis
Saurav Ghosh
Saurav Ghosh
Graduate Research Fellow, Washington University in St. Louis
Machine LearningSystems SecurityBlockchainNetwork Security
A
Abdou Sow
Department of Computer Science and Engineering, Washington University in St. Louis, Missouri, United States
L
Luke Zhang
Department of Computer Science and Engineering, Washington University in St. Louis, Missouri, United States