Distributional Soft Actor-Critic With Three Refinements

📅 2023-10-09
🏛️ IEEE Transactions on Pattern Analysis and Machine Intelligence
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address Q-value overestimation—leading to suboptimal policies—in model-free reinforcement learning, this paper proposes DSAC-T: the first algorithm integrating expectation substitution, twin distributional modeling, and variance-aware gradient adjustment within the Soft Actor-Critic (SAC) framework for dual value distribution learning in continuous action spaces. DSAC-T employs variance-weighted gradient clipping and updates to significantly mitigate training instability induced by stochastic returns and sensitivity to reward scaling. Empirical evaluation across multiple standard benchmarks demonstrates that DSAC-T outperforms SAC, TD3, and DDPG without hyperparameter tuning; it exhibits enhanced training stability, strong robustness to reward scaling, and successful deployment on a real-world wheeled robot control task.
📝 Abstract
Reinforcement learning (RL) has shown remarkable success in solving complex decision-making and control tasks. However, many model-free RL algorithms experience performance degradation due to inaccurate value estimation, particularly the overestimation of Q-values, which can lead to suboptimal policies. To address this issue, we previously proposed the Distributional Soft Actor-Critic (DSAC or DSACv1), an off-policy RL algorithm that enhances value estimation accuracy by learning a continuous Gaussian value distribution. Despite its effectiveness, DSACv1 faces challenges such as training instability and sensitivity to reward scaling, caused by high variance in critic gradients due to return randomness. In this paper, we introduce three key refinements to DSACv1 to overcome these limitations and further improve Q-value estimation accuracy: expected value substitution, twin value distribution learning, and variance-based critic gradient adjustment. The enhanced algorithm, termed DSAC with Three refinements (DSAC-T or DSACv2), is systematically evaluated across a diverse set of benchmark tasks. Without the need for task-specific hyperparameter tuning, DSAC-T consistently matches or outperforms leading model-free RL algorithms, including SAC, TD3, DDPG, TRPO, and PPO, in all tested environments. Additionally, DSAC-T ensures a stable learning process and maintains robust performance across varying reward scales. Its effectiveness is further demonstrated through real-world application in controlling a wheeled robot, highlighting its potential for deployment in practical robotic tasks.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Value Estimation Accuracy
Reward Uncertainty
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
DSAC-T Algorithm
Dynamic Learning Rate Adjustment
🔎 Similar Papers
No similar papers found.
University of Science and Technology Beijing | Tsinghua University
Jingliang Duan
Jingliang Duan
University of Science and Technology Beijing
W
Wenxuan Wang
School of Vehicle and Mobility, Tsinghua University, Beijing, China, 100084
L
Liming Xiao
School of Mechanical Engineering, University of Science and Technology Beijing, China, 100083
J
Jiaxin Gao
School of Vehicle and Mobility, Tsinghua University, Beijing, China, 100084
S
S. Li
School of Vehicle and Mobility, Tsinghua University, Beijing, China, 100084
C
Chang Liu
Y
Ya-Qin Zhang
B
B. Cheng
Keqiang Li
Keqiang Li
Department of Automotive Engineering, Tsinghua University
Intelligent VehiclesAdvanced Driver Assistant Systems