Score
Designs and implements algorithms, objective functions, and training procedures that optimize policies defined over whole trajectories (sequences of states and actions) by estimating and maximizing expected trajectory return. Work includes formulating trajectory-level objectives and joint path probabilities, deriving and implementing pathwise/policy-gradient estimators, and supporting variable-length or step-indexed policies.
This work addresses the challenges of discrete path optimization and black-box utility evaluation in trajectory-based optimal experimental design by modeling trajectories as stochastic variables governed by a parameterized Markov policy. This formulation recasts path optimization as a stochastic optimization problem over policy parameters, thereby introducing probabilistic modeling and Markov decision policies into trajectory-based experimental design for the first time. The approach enables effective exploration of the tail regions of the utility function and is applicable to both linear and nonlinear inverse problems. By integrating a static navigation mesh, a parameterized policy, and black-box utility evaluations, the method demonstrates strong performance in canonical parameter identification tasks, highlighting its broad applicability in model-driven optimal experimental design.
This paper unifies three dominant paradigms in optimal control and motion planning—Model Predictive Path Integral (MPPI) control, reinforcement learning (RL), and diffusion models. Methodologically, it leverages gradient optimization over the Gibbs measure to establish rigorous theoretical connections. The contributions are threefold: (i) MPPI is proven equivalent to gradient ascent on the Gibbs energy functional; (ii) under fixed initial states, policy gradient methods reduce exactly to MPPI; and (iii) the reverse sampling update rule of diffusion models coincides identically with the MPPI update. Collectively, these results establish a fundamental mathematical equivalence among the three frameworks. Beyond unification, the analysis reveals shared mechanistic principles underlying generative planning methods and enables the design of robust, efficient, and interpretable unified generative optimal controllers. This work provides a novel paradigm for tightly integrating learning-based and model-based control.
Policy optimization algorithms suffer from poor interpretability and error-prone implementation due to the complexity of Markov decision process (MDP) modeling and inconsistent use of discounted versus average-reward settings. Method: This paper introduces a unified analytical framework that, for the first time, systematically integrates generalized ergodicity theory with perturbation analysis to characterize the steady-state behavior of diverse policy optimization algorithms under both discounted and average-reward criteria. Contribution/Results: The framework clarifies fundamental algorithmic principles, identifies and corrects common implementation pitfalls, and significantly enhances interpretability and robustness. Empirical validation on MDP modeling and linear quadratic regulator (LQR) benchmarks confirms the framework’s ability to capture algorithmic consistency. Quantitative analysis further demonstrates that minor adjustments to key design parameters exert decisive influence on convergence properties and performance.
Conventional policy optimization relies heavily on approximate dynamic programming or reinforcement learning (RL) frameworks, requiring explicit modeling of actions, control signals, and reward functions—limiting generalizability across diverse control and learning tasks. Method: This paper introduces a novel paradigm that embeds parameterized policies directly into autonomous dynamical systems, enabling joint optimization of policy parameters and system dynamics at the continuous-time dynamical level—bypassing RL-specific abstractions. Contribution/Results: The approach unifies behavior cloning, mechanism design, system identification, and state estimation within a single differentiable optimization framework, without requiring reward engineering or action-space specification. Theoretically, its gradient updates are shown to be equivalent to standard policy gradients, natural gradients, and PPO updates. Empirically, it achieves performance on par with state-of-the-art RL methods across diverse control benchmarks and provides a more intrinsic, fully differentiable foundation for optimizing generative AI systems.
This work addresses the high cost and suboptimality of acquiring high-quality demonstration data in imitation learning, particularly the lack of scalable data sources for goal-conditioned control tasks. To overcome this challenge, the authors propose an efficient data generation and augmentation framework that leverages trajectory optimization to automatically produce thousands of near-optimal trajectories within minutes on a standard laptop. By relabeling intermediate states along these trajectories as new goals, the training dataset is expanded by an order of magnitude. A lightweight goal-conditioned policy trained on this augmented dataset—containing fewer than 80,000 parameters—achieves near-optimal performance and high success rates across multiple tasks. Moreover, its inference speed exceeds that of the trajectory optimization solver by over 6,000×, substantially improving generalization and enabling practical deployment on embedded systems.
This work addresses the challenge that trajectory optimization solvers rely heavily on high-quality initial trajectories, yet solving them in isolation often leads to slow convergence and solution instability. Moreover, conventional diffusion-based strategies suffer from error accumulation over long horizons due to minor deviations. To overcome these limitations, this study introduces, for the first time, feedback gains from a trajectory optimizer into diffusion policy training and proposes a first-order gradient-based Sobolev loss function. By jointly supervising both trajectory samples and their derivatives, the method generates highly accurate initial guesses using only a small number of demonstration trajectories. This enables efficient inference with significantly fewer diffusion steps, markedly mitigates long-horizon error accumulation, and reduces subsequent optimization runtime by a factor of 2 to 20.
This study addresses the challenge of sparse terminal rewards in reinforcement learning for LLM agents, which leaves intermediate steps unsupervised. To tackle this issue, the authors propose T2SPO, a method that leverages historical trajectories to derive remaining-distance targets. T2SPO innovatively incorporates a pretrained TabPFN regressor to estimate state values through in-context dynamic updating rather than gradient-based training, thereby providing step-level auxiliary credit assignment for policy optimization. Experimental evaluations on the ALFWorld and WebShop benchmarks demonstrate that T2SPO significantly improves task success rates for both 1.5B and 7B parameter models, consistently outperforming the GRPO baseline.
This work addresses the challenges of satisfying terminal constraints and achieving high sample efficiency in long-horizon stochastic trajectory optimization. The authors propose a stochastic multi-segment shooting method that stitches together short action sequences optimized under local feedback policies. By leveraging trajectory-based Jacobian approximations—without requiring explicit model gradients—the approach significantly enhances convergence to the terminal set and improves sample efficiency. Framed as a novel paradigm for model-based reinforcement learning under black-box dynamics, the method demonstrates superior performance over existing approaches across three nonlinear underactuated tasks, encompassing both analytical and neural network-based dynamics models, thereby validating its effectiveness in meeting terminal constraints and reducing sample complexity.
This work addresses the low sample efficiency of traditional model-free reinforcement learning methods—such as Proximal Policy Optimization (PPO)—which rely on high-variance advantage estimates. The authors propose Analytic Policy Gradients (APG), a method that leverages differentiable environment dynamics to compute exact, end-to-end gradients of policy returns with respect to policy parameters. To mitigate gradient degradation in long-horizon tasks, APG incorporates a segment-wise backpropagation mechanism and combines Monte Carlo estimation with critic-guided bootstrapping for effective gradient guidance. Evaluated on four continuous control benchmarks under identical network architectures and training protocols, APG consistently outperforms PPO, demonstrating substantially higher sample efficiency and faster convergence.
This study addresses the challenge of real-time path planning for autonomous vehicles in environments containing circular no-fly threat zones, where conventional optimal control methods suffer from high computational complexity. To overcome this limitation, the authors propose a reinforcement learning framework based on Deep Deterministic Policy Gradient (DDPG), employing an Actor-Critic architecture and a carefully designed reward function to enable direct mapping from states to actions, thereby rapidly generating safe and feasible trajectories. Notably, the approach innovatively leverages DDPG to construct a “feasibility set” for path planning, offering a priori judgment of task realizability before execution. Simulation results demonstrate that, within this feasibility set, the method achieves significantly higher computational efficiency than pseudospectral optimal control, making it suitable for real-time applications—albeit at the cost of global optimality—while effectively avoiding infeasible regions.