🤖 AI Summary
This study addresses the bottlenecks of slow inference, fixed contact points, and the absence of bidirectional coupling in neural robot dynamics models by proposing a differentiable neurodynamics framework that integrates a parallel streaming architecture with a universal contact encoder. This method introduces the first bidirectional coupling mechanism supporting contacts at arbitrary locations and employs an analytical solver hybrid technique for efficient modeling. Experimental results demonstrate that the proposed framework accelerates dynamics prediction by 55× and PPO policy training by 3.8×, while significantly improving both the efficiency and accuracy of MPC planning. These advances effectively overcome the limitations of conventional models regarding computational speed and scene generalization.
📝 Abstract
Compared with analytical physics, learned dynamics models promise robot simulation that is faster, inherently differentiable, and easily adaptable to real data. Neural Robot Dynamics (NeRD) pursues this by keeping collision detection analytical and replacing a simulator's numerical dynamics for the robot with a learned model. Three limitations remain. NeRD offers little speedup over the simulator it learned from, accepts contact only at predefined points, and has no two-way coupling with objects it manipulates. FlashNeRD removes all three with a parallel streaming architecture that makes each prediction faster and more accurate, an encoder that accepts contacts wherever they occur, and two-way coupling with objects simulated by analytical solvers. Experiments across five robots show faster and more accurate dynamics, faster policy learning, and faster inference-time planning. Across three robots, FlashNeRD's dynamics model is up to $55\times$ faster than an optimized NeRD and more accurate over long rollouts. This speedup extends to policy learning, where PPO trains an ANYmal locomotion policy in 34 seconds, $3.8\times$ faster than the analytical simulator and $2.7\times$ faster than an optimized NeRD. With DIAL-MPC, a sampling-based MPC method, the robot climbs all six test platforms where fixed-contact NeRD manages one, at approximately half the analytical simulator's planning time. Cube-reorientation policies trained with FlashNeRD complete within 2% of the simulator-trained policy's target count.