ADP: Adversarial Dynamics Priors for Physically Grounded Humanoid Locomotion

📅 2026-07-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing motion-prior-based methods model only kinematic features while neglecting dynamic characteristics—such as center-of-mass dynamics, momentum, and contact forces—resulting in insufficient disturbance rejection for humanoid robots. This work proposes Adversarial Dynamics Priors (ADP), which extends adversarial imitation learning from the kinematic to the dynamic domain for the first time. ADP constructs a reference distribution centered on centroidal dynamics through trajectory optimization and employs a temporal window discriminator to guide policy learning within a reinforcement learning framework that does not require explicit action tracking, thereby generating dynamically consistent and robust motions. Compared to the strongest baseline, AMP, ADP improves the impact threshold at 80% success rate by 16.7%, reduces average recovery time by 47.9%, and decreases velocity tracking error by 35.4%.
📝 Abstract
In this paper, we propose Adversarial Dynamics Priors (ADP) for perturbation-resilient humanoid locomotion control. Existing motion prior-based methods induce natural motion styles by imitating kinematic motion features, but they do not directly regularize dynamics features, such as CoM motion, centroidal momentum, contact forces, and contact states. To address this limitation, we replace kinematic motion-style feature with selected dynamics features extracted from locomotion trajectories as the target of adversarial regularization.To this end, we use trajectory optimization to construct a reference dataset and train a discriminator to evaluate whether policy-induced temporal windows are consistent with the resulting reference distribution.Without explicit motion tracking, ADP encourages policy rollouts to remain close to the reference support, even after perturbations. Experimental results show that, compared with AMP, the strongest baseline in our evaluation, ADP improves the $80\%$-success impulse threshold ($J_{80}$) by $16.7\%$, while reducing direction-averaged recovery time and velocity tracking error by $47.9\%$ and $35.4\%$, respectively.
Problem

Research questions and friction points this paper is trying to address.

humanoid locomotion
motion priors
dynamics regularization
perturbation resilience
adversarial learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial Dynamics Priors
humanoid locomotion
dynamics regularization
perturbation resilience
trajectory optimization
🔎 Similar Papers
No similar papers found.
Seokju Lee
Seokju Lee
Ph.D. Student, KAIST MSC Lab
legged roboticslearningcontrol
J
Jeongtae Lee
Mechatronics, Systems and Control Lab (MSC Lab), Department of Mechanical Engineering, Korea Advanced Institute of Science and Technology (KAIST), Yuseong-gu, Daejeon 34141, Republic of Korea
J
Jeonghyeok Lim
Mechatronics, Systems and Control Lab (MSC Lab), Department of Mechanical Engineering, Korea Advanced Institute of Science and Technology (KAIST), Yuseong-gu, Daejeon 34141, Republic of Korea
J
Jeonguk Kang
Samsung Electronics, Future Robotics AI Group, Seoul, Republic of Korea
B
Byungwook Lee
Mechatronics, Systems and Control Lab (MSC Lab), Department of Mechanical Engineering, Korea Advanced Institute of Science and Technology (KAIST), Yuseong-gu, Daejeon 34141, Republic of Korea
S
Seungho Han
School of Electrical Engineering, Hanyang University, Ansan 15588, Republic of Korea
Keun Ha Choi
Keun Ha Choi
한국과학기술원 연구교수
자율주행
D
Dongil Park
Advanced Robotics Research Center, Korea Institute of Machinery & Materials (KIMM), Daejeon 34103, Republic of Korea
Kyung-Soo Kim
Kyung-Soo Kim
Professor of Mechanical Engineering, KAIST
controlrobotmechatronicsmanufacturingautomation