🤖 AI Summary
This work addresses the common issue in hybrid discrete–continuous planning where first-order trajectories generated by conventional methods often violate second-order dynamical constraints of robotic systems, rendering them infeasible for execution. To bridge the gap between high-level task planning and low-level physical execution, the authors propose a reinforcement learning–based trajectory refinement framework that explicitly embeds analytical second-order dynamics into a Markov decision process. This approach continuously optimizes first-order trajectories produced by a high-level hybrid planner while respecting constraints on time windows, velocity, and acceleration. By integrating reinforcement learning with explicit second-order dynamical modeling—a combination not previously explored—the method significantly enhances the physical feasibility and real-world executability of planned trajectories.
📝 Abstract
In many robotic tasks, agents must traverse a sequence of spatial regions to complete a mission. Such problems are inherently mixed discrete-continuous: a high-level action sequence and a physically feasible continuous trajectory. The resulting trajectory and action sequence must also satisfy problem constraints such as deadlines, time windows, and velocity or acceleration limits. While hybrid temporal planners attempt to address this challenge, they typically model motion using linear (first-order) dynamics, which cannot guarantee that the resulting plan respects the robot's true physical constraints. Consequently, even when the high-level action sequence is fixed, producing a dynamically feasible trajectory becomes a bi-level optimization problem. We address this problem via reinforcement learning in continuous space. We define a Markov Decision Process that explicitly incorporates analytical second-order constraints and use it to refine first-order plans generated by a hybrid planner. Our results show that this approach can reliably recover physical feasibility and effectively bridge the gap between a planner's initial first-order trajectory and the dynamics required for real execution.