Score
Design and implement policy- and trajectory-optimization algorithms that integrate trajectory-derived guidance (prompts, spatial priors, or correction signals) into reinforcement-learning updates; build RL-corrected optimizers and analyze their effects on trajectory feasibility, reliability, and convergence of the optimization process.
This paper addresses the challenge of deeply integrating model predictive control (MPC) and reinforcement learning (RL), stemming from their fundamentally divergent model usage paradigms. To resolve this, we propose the first unified taxonomy for MPC–RL fusion, centered on *how models are used*, categorizing approaches into three paradigms: MPC-augmented RL, RL-augmented MPC, and co-designed architectures. Leveraging a unified Actor–Critic modeling framework, we systematically analyze how MPC’s online optimization enhances RL’s closed-loop performance and establish a performance-gain-oriented evaluation perspective grounded in closed-loop metrics. The survey comprehensively covers six application domains—including robotics, energy systems, and autonomous driving—and synthesizes cross-cutting modeling techniques bridging control theory and RL. Our work provides a scalable methodology and principled design guidelines for hybrid intelligent control systems.
To address the challenge of real-time flight path replanning for civil aviation under emergency conditions, this paper proposes a reinforcement learning (RL)-guided hybrid path optimization method. The approach leverages an RL agent to pre-generate high-quality candidate paths, which dynamically constrain the search space of classical planning algorithms—such as A* or RRT—thereby significantly improving computational efficiency while preserving solution quality. The core contribution lies in the synergistic integration of data-driven learning and model-based search, enabling verifiable, deployable trajectory planning. Experimental results demonstrate that the generated paths incur less than 1% fuel consumption deviation from globally optimal solutions, while achieving up to a 50% reduction in replanning latency compared to conventional methods. Unlike purely learning-based or purely search-based approaches, the proposed framework achieves a balanced trade-off between real-time responsiveness and optimality, establishing a novel, rigorous paradigm for airborne emergency decision-making.
This work investigates under what conditions large language models (LLMs) can serve as general-purpose black-box policy optimizers in place of conventional reinforcement learning algorithms. To this end, the authors propose Prompted Policy Optimization (PromptPO), a method that iteratively generates executable policies by prompting an LLM with Python-based descriptions of the state space, action space, and reward function, and refines them using environmental feedback. The first systematic evaluation demonstrates that PromptPO automatically synthesizes diverse policies—ranging from rule-based controllers to planning algorithms—and matches or exceeds standard RL baselines with fewer environment interactions on challenging exploration tasks, Meta-World benchmarks, and several real-world control problems. However, it remains limited in domains requiring fine-grained continuous control, such as MuJoCo tasks.
This work addresses the inefficiency of existing on-policy reinforcement learning exploration methods, which often overlook the intrinsic value of states and struggle to discover high-reward trajectories. The authors propose a directed exploration mechanism grounded in a differentiable dynamics model, integrating task objectives and physics-informed guidance into the exploration process through analytical policy gradients. This approach represents the first use of analytical policy gradients to drive purposeful exploration, departing from conventional paradigms that rely indiscriminately on entropy maximization or state novelty. Empirical results demonstrate that the proposed framework significantly accelerates policy convergence and enhances final performance on robotic control tasks, exhibiting superior efficiency and stability in exploration and learning.
Addressing practical challenges in reinforcement learning—such as sparse and delayed rewards and training instability—this paper presents a systematic survey of reward engineering and reward shaping. We propose the first fine-grained taxonomy of reward design techniques, explicitly exposing their implicit assumptions and failure boundaries. Furthermore, we introduce an evaluation framework for reward shaping that jointly balances interpretability and empirical effectiveness. Our analysis integrates theoretical foundations of RL, deep RL practice, formal modeling of reward functions, and cross-domain applications—including robotics and autonomous driving. This work fills a critical gap by providing the first comprehensive, methodology-driven survey of reward design. It establishes a unified tripartite research framework comprising methodology, taxonomic classification, and application boundaries. The resulting synthesis delivers a reproducible, transferable engineering guide for algorithm designers, significantly enhancing the robustness and real-world deployability of RL systems. (149 words)
This work addresses the limited interpretability of internal mechanisms in reinforcement learning—such as value estimation, policy optimization, and their interaction with temporal difference (TD) signals—by proposing the first systematic, multi-perspective framework for visualizing loss landscapes. The approach integrates the geometric structure of value functions, policy optimization trajectories, TD error dynamics, and state-driven regions through techniques including 3D loss surface reconstruction, policy landscape visualization under a frozen critic, joint trajectories of time–Bellman error–policy weights, and state–TD mappings. Applied to the ADHDP algorithm for spacecraft attitude control, the framework enables comparative analysis of multiple variants, revealing how training stabilizers and target update mechanisms reshape the optimization landscape and influence learning stability. This study establishes a new paradigm for interpretable reinforcement learning and provides actionable insights for algorithm design.
This work addresses the challenges of data scarcity and suboptimal quality in offline reinforcement learning, which arise from reliance on a limited number of imperfect trajectories. To mitigate these issues, the paper proposes a trajectory-level data augmentation method that leverages the geometric relationships among the reward function, value function, and behavior policy. By incorporating the intrinsic geometry of the task, the approach remains compatible with suboptimal data-collecting policies and provides theoretical justification for trajectory augmentation. This is the first study to integrate trajectory-level augmentation with task-specific geometric structure, demonstrating significant improvements in both performance and data efficiency across a range of high-dimensional and partially observable navigation tasks.
This work addresses the lack of a unified theoretical framework for reinforcement learning, which has hindered systematic analysis of its convergence, sample complexity, and generalization. Building upon Markov decision processes and Bellman operators, the paper introduces a cohesive analytical framework that integrates tools from operator theory, stochastic approximation, convex duality, and function approximation. This framework encompasses a broad range of algorithms, including value iteration, policy iteration, temporal difference methods, off-policy learning, and constrained MDPs. By leveraging contraction mappings, monotone operators, martingale techniques, mirror/proximal optimization, concentration inequalities, and mixing process theory, the study establishes finite-sample performance bounds and asymptotic convergence guarantees for diverse reinforcement learning algorithms, thereby forging a rigorous theoretical bridge between probability theory, optimization, and statistics.
This work addresses the longstanding methodological, objective, and cultural divide between reinforcement learning and control theory by proposing a novel paradigm that integrates adaptive control with actor-critic reinforcement learning. The resulting framework enables data-driven optimization of controllers by unifying dynamic programming and online learning mechanisms, thereby reconciling modeling and optimization perspectives from both fields within classical motion control tasks. Theoretical analysis elucidates fundamental differences between the two approaches, while empirical results demonstrate the efficacy of the integrated strategy. This synthesis offers a solution for controlling systems with unknown dynamics that simultaneously guarantees stability and retains strong learning capabilities, fostering interoperability and synergistic development across disciplinary boundaries.
This work addresses key limitations of existing reinforcement learning (RL) approaches for autonomous driving—namely, poor interpretability, difficulty in modeling road geometry, and weak compatibility with modern planning architectures—by proposing a novel RL framework that integrates the Frenet coordinate system with polynomial trajectory planning. The method introduces a kinematic feasibility check after policy inference to generate trajectories that respect vehicle dynamic constraints. To the best of our knowledge, this is the first approach to embed Frenet coordinates and analytical trajectory planning within an end-to-end RL pipeline, substantially enhancing interpretability and reducing learning complexity. Evaluated on the CARLA Offline Leaderboard v1 and NoCrash benchmarks, the proposed method achieves a 5% and 11% improvement in driving score, and an 8% and 19% increase in task success rate, respectively, significantly outperforming current control-oriented RL baselines.