ADP: Adversarial Dynamics Priors for Physically Grounded Humanoid Locomotion
2026-07-03 • Robotics
RoboticsMachine Learning
AI summaryⓘ
The authors present a new method called Adversarial Dynamics Priors (ADP) to help humanoid robots walk better even when pushed or disturbed. Unlike earlier methods that focused on copying how the robot looks while moving, ADP focuses on the robot's physical balance and forces during walking. They train a system to recognize good walking patterns based on physical dynamics, encouraging the robot to stick to these patterns even after being pushed. Their tests show that ADP helps the robot recover faster and track its movements more accurately compared to the best previous method.
Humanoid locomotionMotion priorsDynamics featuresTrajectory optimizationAdversarial regularizationCentroidal momentumContact forcesImpulse thresholdPolicy rolloutsRecovery time
Authors
Seokju Lee, Jeongtae Lee, Jeonghyeok Lim, Jeonguk Kang, Byungwook Lee, Seungho Han, Keun Ha Choi, Dongil Park, Kyung-Soo Kim
Abstract
In this paper, we propose Adversarial Dynamics Priors (ADP) for perturbation-resilient humanoid locomotion control. Existing motion prior-based methods induce natural motion styles by imitating kinematic motion features, but they do not directly regularize dynamics features, such as CoM motion, centroidal momentum, contact forces, and contact states. To address this limitation, we replace kinematic motion-style feature with selected dynamics features extracted from locomotion trajectories as the target of adversarial regularization.To this end, we use trajectory optimization to construct a reference dataset and train a discriminator to evaluate whether policy-induced temporal windows are consistent with the resulting reference distribution.Without explicit motion tracking, ADP encourages policy rollouts to remain close to the reference support, even after perturbations. Experimental results show that, compared with AMP, the strongest baseline in our evaluation, ADP improves the $80\%$-success impulse threshold ($J_{80}$) by $16.7\%$, while reducing direction-averaged recovery time and velocity tracking error by $47.9\%$ and $35.4\%$, respectively.