๐ค AI Summary
Deterministic policy gradient methods (e.g., DDPG, TD3) often suffer from local optima in complex tasksโsuch as dexterous manipulation and constrained motion controlโdue to multimodality in the Q-function, which misguides actor updates. To address this, we propose a Multi-Actor Collaborative Optimization framework with a Differentiable Proxy Q-function. Our approach jointly optimizes multiple deterministic actors to generate diverse action candidates, employs a learnable, smooth proxy Q-network to provide globally consistent gradient signals, integrates offline policy optimization, and applies action-space reparameterization to overcome limitations of single-point gradient ascent. Evaluated on dexterous manipulation, constrained motion control, and large-scale discrete recommendation tasks, our method achieves a 37% improvement in optimal action discovery rate over DDPG, TD3, and other baselines, demonstrating superior exploration capability and convergence robustness in high-dimensional, non-convex policy optimization landscapes.
๐ Abstract
In reinforcement learning, off-policy actor-critic approaches like DDPG and TD3 are based on the deterministic policy gradient. Herein, the Q-function is trained from off-policy environment data and the actor (policy) is trained to maximize the Q-function via gradient ascent. We observe that in complex tasks like dexterous manipulation and restricted locomotion, the Q-value is a complex function of action, having several local optima or discontinuities. This poses a challenge for gradient ascent to traverse and makes the actor prone to get stuck at local optima. To address this, we introduce a new actor architecture that combines two simple insights: (i) use multiple actors and evaluate the Q-value maximizing action, and (ii) learn surrogates to the Q-function that are simpler to optimize with gradient-based methods. We evaluate tasks such as restricted locomotion, dexterous manipulation, and large discrete-action space recommender systems and show that our actor finds optimal actions more frequently and outperforms alternate actor architectures.