Mitigating Suboptimality of Deterministic Policy Gradients in Complex Q-functions

๐Ÿ“… 2024-10-15
๐Ÿ›๏ธ arXiv.org
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Deterministic policy gradient methods (e.g., DDPG, TD3) often suffer from local optima in complex tasksโ€”such as dexterous manipulation and constrained motion controlโ€”due to multimodality in the Q-function, which misguides actor updates. To address this, we propose a Multi-Actor Collaborative Optimization framework with a Differentiable Proxy Q-function. Our approach jointly optimizes multiple deterministic actors to generate diverse action candidates, employs a learnable, smooth proxy Q-network to provide globally consistent gradient signals, integrates offline policy optimization, and applies action-space reparameterization to overcome limitations of single-point gradient ascent. Evaluated on dexterous manipulation, constrained motion control, and large-scale discrete recommendation tasks, our method achieves a 37% improvement in optimal action discovery rate over DDPG, TD3, and other baselines, demonstrating superior exploration capability and convergence robustness in high-dimensional, non-convex policy optimization landscapes.

Technology Category

Search and Optimization: Learning to SearchIntelligent Robots: Learning & Optimization for ROBMachine Learning: Optimization

Application Category

Economics, Online Markets and Human Computation: Economics and fairness of platforms and recommendation systemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalization
๐Ÿ“ Abstract
In reinforcement learning, off-policy actor-critic approaches like DDPG and TD3 are based on the deterministic policy gradient. Herein, the Q-function is trained from off-policy environment data and the actor (policy) is trained to maximize the Q-function via gradient ascent. We observe that in complex tasks like dexterous manipulation and restricted locomotion, the Q-value is a complex function of action, having several local optima or discontinuities. This poses a challenge for gradient ascent to traverse and makes the actor prone to get stuck at local optima. To address this, we introduce a new actor architecture that combines two simple insights: (i) use multiple actors and evaluate the Q-value maximizing action, and (ii) learn surrogates to the Q-function that are simpler to optimize with gradient-based methods. We evaluate tasks such as restricted locomotion, dexterous manipulation, and large discrete-action space recommender systems and show that our actor finds optimal actions more frequently and outperforms alternate actor architectures.
Problem

Research questions and friction points this paper is trying to address.

Addresses deterministic policy gradients getting stuck in local optima
Improves Q-function optimization in complex reinforcement learning tasks
Enhances action selection in constrained locomotion and manipulation tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generates multiple action proposals for selection
Approximates Q-function by truncating poor local optima
Uses actor architecture to guide gradient ascent effectively
University of Southern California | Line Yahoo Corp | NAVER Cloud | KAIST
A
Ayush Jain
University of Southern California
Norio Kosaka
Norio Kosaka
Line Yahoo Corp
X
Xinhu Li
University of Southern California
K
Kyung-Min Kim
NAVER Cloud
E
Erdem Biyik
University of Southern California
J
Joseph J Lim
KAIST