Score
Design and implement neural-network–based policies and value-function critics for sequential decision-making with continuous or high-dimensional states and actions, and build actor–critic architectures and deterministic policy-gradient algorithms (e.g., DDPG) to learn control policies. Analyze and tune training procedures — including off-policy learning, exploration strategies, reward shaping and constraint handling — to improve sample efficiency, stability, and progressive performance gains when optimizing nonconvex objectives through interaction.
In discrete-action offline reinforcement learning, actor-critic methods (e.g., SAC) underperform value-based approaches (e.g., DQN), primarily due to entropy coupling between the actor and critic. Method: This paper identifies such coupling as the key bottleneck and proposes a novel entropy decoupling mechanism—separately regularizing actor and critic entropies. Building upon this, we introduce a general offline actor-critic framework supporting *m*-step Bellman updates, flexible policy optimization schemes, and theoretically guaranteed convergence. The framework unifies standard policy gradients and entropy-regularized objectives, and also accommodates exploration-free, entropy-unregularized training. Contribution/Results: Evaluated on the Atari benchmark, our method achieves performance on par with DQN—without entropy regularization or auxiliary exploration mechanisms—thereby substantially advancing both the practical applicability and theoretical soundness of offline policy learning in discrete action spaces.
This work addresses the challenges of credit assignment and high gradient variance in reinforcement learning with hybrid discrete-continuous action spaces, where conventional policy gradient methods suffer significant performance degradation, particularly in high-dimensional continuous action settings. To overcome these limitations, the paper proposes Hybrid Policy Optimization (HPO), which integrates pathwise derivatives and score function gradients to construct an unbiased hybrid gradient estimator within differentiable simulators. Theoretical analysis reveals that the cross-term in the hybrid gradient becomes negligible near discrete optimal responses, justifying an approximately decoupled update strategy that effectively reduces variance. Empirical results demonstrate that HPO substantially outperforms Proximal Policy Optimization (PPO) on inventory control and switched linear quadratic regulator tasks, with performance gains increasing as the dimensionality of the continuous action space grows.
Deterministic policy gradient methods (e.g., DDPG, TD3) often suffer from local optima in complex tasks—such as dexterous manipulation and constrained motion control—due to multimodality in the Q-function, which misguides actor updates. To address this, we propose a Multi-Actor Collaborative Optimization framework with a Differentiable Proxy Q-function. Our approach jointly optimizes multiple deterministic actors to generate diverse action candidates, employs a learnable, smooth proxy Q-network to provide globally consistent gradient signals, integrates offline policy optimization, and applies action-space reparameterization to overcome limitations of single-point gradient ascent. Evaluated on dexterous manipulation, constrained motion control, and large-scale discrete recommendation tasks, our method achieves a 37% improvement in optimal action discovery rate over DDPG, TD3, and other baselines, demonstrating superior exploration capability and convergence robustness in high-dimensional, non-convex policy optimization landscapes.
Deterministic Policy Gradient (DPG) algorithms rely on precise gradient estimates of the action-value function provided by the critic; however, under function approximation, such gradients are prone to bias and estimation error, leading to unstable policy updates. To address this, we propose Zeroth-Order Deterministic Policy Gradient (ZODPG), the first method to introduce two-point stochastic gradient estimation directly in action space—thereby eliminating explicit differentiation of the action-value function and theoretically ensuring compatibility of policy updates. By integrating zeroth-order optimization into the actor-critic framework, ZODPG avoids gradient approximation errors inherent in first-order DPG variants. Empirical evaluation across multiple continuous-control benchmark tasks demonstrates that ZODPG significantly reduces gradient estimation error, achieves more robust convergence, and outperforms existing state-of-the-art methods in both sample efficiency and final performance.
To address insufficient exploration and inefficient reward exploitation in sparse-reward continuous-control tasks under DDPG, this paper proposes three synergistic improvements: (1) a time-varying εₜ-greedy policy to enhance state-space coverage; (2) a dual experience replay buffer (GDRB) that separates high- and low-return trajectories to improve sample discriminability; and (3) longest-n-step return estimation to strengthen temporal propagation of sparse positive rewards. All modifications require no additional networks or model assumptions, thereby preserving algorithmic simplicity while significantly improving training stability and convergence speed. Evaluated on standard sparse-reward benchmarks—including AntMaze and Sparse HalfCheetah—the method outperforms the original DDPG as well as state-of-the-art approaches (TD3, SAC-Sparse). Ablation studies confirm that each component contributes significantly and complementarily to overall performance.
This work addresses the issue of undefined and unstable gradients in deterministic policy gradient methods under sparse or discrete reward settings, where the Q-function is non-differentiable with respect to actions. To overcome this limitation, the paper proposes Soft Deterministic Policy Gradient (Soft-DPG), which introduces Gaussian smoothing into the deterministic policy gradient framework for the first time. By constructing a smoothed Bellman equation and redefining the action-value function, Soft-DPG circumvents the explicit reliance on the gradient of the Q-function with respect to actions. Theoretical analysis demonstrates that the proposed method ensures well-defined policy gradients even when the Q-function is non-smooth. Empirical results show that Soft-DPG achieves competitive performance in standard continuous control tasks with dense rewards and significantly outperforms DDPG in environments with sparse or discrete rewards.
This study addresses the challenge of real-time path planning for autonomous vehicles in environments containing circular no-fly threat zones, where conventional optimal control methods suffer from high computational complexity. To overcome this limitation, the authors propose a reinforcement learning framework based on Deep Deterministic Policy Gradient (DDPG), employing an Actor-Critic architecture and a carefully designed reward function to enable direct mapping from states to actions, thereby rapidly generating safe and feasible trajectories. Notably, the approach innovatively leverages DDPG to construct a “feasibility set” for path planning, offering a priori judgment of task realizability before execution. Simulation results demonstrate that, within this feasibility set, the method achieves significantly higher computational efficiency than pseudospectral optimal control, making it suitable for real-time applications—albeit at the cost of global optimality—while effectively avoiding infeasible regions.
This work addresses the computational burden and numerical instability of flow-matching–based reinforcement learning in real-world robotic manipulation, which stems from its reliance on backpropagation through time (BPTT). To overcome this limitation, the authors propose FlowDPG, a BPTT-free DDPG-style algorithm that distills critic gradients into a velocity field, thereby integrating demonstration-driven motion with critic-guided corrections during policy optimization. Theoretical analysis demonstrates that FlowDPG’s update direction aligns with the classical deterministic policy gradient. Evaluated on a multi-stage dual-arm AirPods assembly task, FlowDPG achieves a 92% end-to-end success rate, substantially outperforming existing approaches based on value conditioning, auxiliary module adaptation, and adjoint gradients.
This work addresses the low sample efficiency of traditional model-free reinforcement learning methods—such as Proximal Policy Optimization (PPO)—which rely on high-variance advantage estimates. The authors propose Analytic Policy Gradients (APG), a method that leverages differentiable environment dynamics to compute exact, end-to-end gradients of policy returns with respect to policy parameters. To mitigate gradient degradation in long-horizon tasks, APG incorporates a segment-wise backpropagation mechanism and combines Monte Carlo estimation with critic-guided bootstrapping for effective gradient guidance. Evaluated on four continuous control benchmarks under identical network architectures and training protocols, APG consistently outperforms PPO, demonstrating substantially higher sample efficiency and faster convergence.
This work addresses the challenge of modeling long-term dependencies in non-Markovian decision processes, where observations and rewards depend on the full interaction history, rendering conventional reinforcement learning methods ineffective. The authors propose a reward-oriented Agent State Markov (ASM) policy framework that recursively updates an internal state representation and jointly optimizes this representation with the control policy. Within this framework, they establish the first policy gradient theorem applicable to non-Markovian environments. Building on this theoretical foundation, they develop an end-to-end ASMPG algorithm and provide rigorous proofs of its finite-time and almost sure convergence. Empirical evaluations demonstrate that the proposed method significantly outperforms baseline approaches relying on predictive objectives for state representation across multiple non-Markovian tasks, confirming the effectiveness and superiority of the ASM paradigm in handling history-dependent dynamics.