๐ค AI Summary
This study addresses the risk of hardware damage in humanoid robots executing highly dynamic motions due to policy failures or external disturbances. We propose VAPS (Viability-based Adaptive Protective Selection), a framework that formulates safety as a receding-horizon decision-making problem. By employing a learned viability predictor to evaluate, in real time, the execution conditions of nominal, abort, and protective falling policies, VAPS constructs a hierarchical decision mechanism that dynamically selects the minimum-cost strategy. Experiments on physical Unitree G1 and LimX Oli platforms demonstrate that VAPS significantly reduces head and hand ground-contact frequencies. The proposed approach achieves a Pareto-optimal balance between task success rate and safety, effectively ensuring hardware integrity during extreme maneuvers.
๐ Abstract
Dynamic humanoid motions such as flips risk hardware damage due to suboptimal policies, disturbances or sim-to-real gaps. A motion tracking policy offers no way out once the maneuver leaves its reference, and a backup policy needs to take over to protect the hardware for a minimum-damage landing. Which backup to use matters as much as when to switch. We present Viability-Aware Policy Selection (VAPS), which treats safety as a policy-conditioned, receding-horizon decision. Besides a protective fall policy, we also train an abort policy which can abort the motion at any time, landing on its feet. At every control step, learned predictors estimate whether the nominal tracking policy and the abort policy remain viable over a short horizon, and a least-sacrificial hierarchy keeps the most task-ambitious behavior that remains viable. In simulation with randomized disturbances, VAPS sharply reduces head contact and hand contact, which are the dominant sources of hardware damage, with both a Unitree G1 and a LimX Oli; on the LimX Oli, we validate the viability predictors and the full VAPS controller for side-flip motions. VAPS Pareto-dominates the strongest single-network alternatives we could train, including an end-to-end safe-tracking policy and students distilled from VAPS's own oracle-routed decisions, in both task success and head impact. We also show that VAPS is a powerful framework to supervise undertrained policies and protect the hardware.