🤖 AI Summary
This study addresses the susceptibility of trajectory optimization for nonlinear dynamics to local optima and its limited exploration capabilities by proposing the PER-DDP framework. Grounded in the free energy inequality, this method constructs an entropy-regularized population-based framework that integrates prior-guided sampling, extended evaluation, and Pareto filtering mechanisms. These components effectively decouple the sampling process from population size while preserving the second-order convergence properties of differential dynamic programming and significantly enhancing global search capability. Experimental evaluations across multiple dynamical systems and hundreds of environments demonstrate that the proposed framework consistently outperforms existing state-of-the-art methods in success rate. Furthermore, PER-DDP reliably discovers solutions that remain inaccessible to baseline algorithms, highlighting its superior exploration capacity and practical effectiveness for complex trajectory optimization tasks.
📝 Abstract
Trajectory optimization (TO) under nonlinear dynamics, actuation limits and collision avoidance constraints is a fundamental problem in robotics, albeit especially challenging due to its highly non-convex nature. For this setting, Differential Dynamic Programming (DDP) is an efficient second-order shooting method, yet its local structure renders it vulnerable to suboptimal basins. Sampling-augmented variants mitigate this susceptibility through stochastic exploration, but often sample only around the few trajectories they retain for reoptimization, based solely on their cost which restricts exploration breadth. We introduce Pareto-Optimal Entropy-Regularized DDP (PER-DDP), an entropy-regularized population framework derived from the free-energy/relative-entropy inequality. Our method combines prior-guided sampling that shapes exploration around each retained trajectory, with expanded rollout evaluations that probe these sampling policies beyond the few retained candidates, and Pareto filtering for preserving task-constraint alternatives across iterations. This decouples sampling effort from the optimization population size and broadens exploration without sacrificing the second-order structure that makes DDP effective. Across multiple systems and hundreds of environments, PER-DDP achieves higher success rates than state-of-the-art sampling-augmented TO methods and finds reliable solutions in environments beyond the reach of all baselines.