Pareto-Optimal Entropy-Regularized Trajectory Optimization

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the susceptibility of trajectory optimization for nonlinear dynamics to local optima and its limited exploration capabilities by proposing the PER-DDP framework. Grounded in the free energy inequality, this method constructs an entropy-regularized population-based framework that integrates prior-guided sampling, extended evaluation, and Pareto filtering mechanisms. These components effectively decouple the sampling process from population size while preserving the second-order convergence properties of differential dynamic programming and significantly enhancing global search capability. Experimental evaluations across multiple dynamical systems and hundreds of environments demonstrate that the proposed framework consistently outperforms existing state-of-the-art methods in success rate. Furthermore, PER-DDP reliably discovers solutions that remain inaccessible to baseline algorithms, highlighting its superior exploration capacity and practical effectiveness for complex trajectory optimization tasks.
📝 Abstract
Trajectory optimization (TO) under nonlinear dynamics, actuation limits and collision avoidance constraints is a fundamental problem in robotics, albeit especially challenging due to its highly non-convex nature. For this setting, Differential Dynamic Programming (DDP) is an efficient second-order shooting method, yet its local structure renders it vulnerable to suboptimal basins. Sampling-augmented variants mitigate this susceptibility through stochastic exploration, but often sample only around the few trajectories they retain for reoptimization, based solely on their cost which restricts exploration breadth. We introduce Pareto-Optimal Entropy-Regularized DDP (PER-DDP), an entropy-regularized population framework derived from the free-energy/relative-entropy inequality. Our method combines prior-guided sampling that shapes exploration around each retained trajectory, with expanded rollout evaluations that probe these sampling policies beyond the few retained candidates, and Pareto filtering for preserving task-constraint alternatives across iterations. This decouples sampling effort from the optimization population size and broadens exploration without sacrificing the second-order structure that makes DDP effective. Across multiple systems and hundreds of environments, PER-DDP achieves higher success rates than state-of-the-art sampling-augmented TO methods and finds reliable solutions in environments beyond the reach of all baselines.
Problem

Research questions and friction points this paper is trying to address.

Trajectory Optimization
Differential Dynamic Programming
Non-convex Optimization
Sampling-based Exploration
Robotics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Entropy-Regularized DDP
Trajectory Optimization
Pareto Filtering
Prior-Guided Sampling
Differential Dynamic Programming
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Dimitrios S. Georgiou
Daniel Guggenheim School of Aerospace Engineering, Georgia Institute of Technology, Atlanta, GA, USA
Augustinos D. Saravanos
Augustinos D. Saravanos
Postdoctoral Researcher, Massachusetts Institute of Technology
OptimizationMachine LearningControl TheoryMulti-Agent SystemsLarge-Scale Decision-Making
E
Evangelos A. Theodorou
Daniel Guggenheim School of Aerospace Engineering, Georgia Institute of Technology, Atlanta, GA, USA