π€ AI Summary
This study addresses the limitation of traditional Actor-Critic methods, which treat historical data merely as off-policy samples without leveraging cumulative feedback for policy optimization. We propose RDA2C, a method that assigns replay data a novel role as an "empirical dual objective." By fitting an advantage regression model via regularized dual averaging and integrating entropy mirror mapping, Generalized Advantage Estimation (GAE), and Twin-Q networks, RDA2C generates the current policy such that historical advantage estimates directly contribute to the actor objective. Furthermore, we establish a theoretical framework for finite-time value gap decomposition. Extensive evaluations on MuJoCo and Atari benchmarks demonstrate that RDA2C surpasses baselines such as PPO, achieves performance comparable to SAC, and significantly improves sample efficiency.
π Abstract
Actor-critic methods reuse past experience to improve sample efficiency. However, historical data are typically regarded as off-policy samples for the current policy-improvement update. This work introduces Regularized Dual Averaging Actor Critic (RDA2C), which assigns a distinct role to replay. In regularized dual averaging, the subsequent policy is determined by accumulated policy-improvement feedback, so historical advantage estimates contribute directly to the actor objective rather than solely to the most recent update. RDA2C stores state-action samples with critic-estimated advantage labels, fits a dual score model $Z_\theta$ to the aggregated dataset, and derives the current policy from the accumulated score model using the entropy mirror map. In this way, replay defines an empirical dual objective from which the policy is computed. To analyze RDA2C, we establish a finite-time value-gap decomposition, separating the regularized dual-averaging term from errors due to stale-replay supervised fitting, critic bias, finite-buffer variance, and replay coverage, and stating the assumptions under which each error is bounded. RDA2C accepts advantage labels from any critic. With GAE labels, RDA2C outperforms PPO on six of eight MuJoCo tasks and eight of twelve Atari games. RDA2C also outperforms AAPDA, the closest dual-averaging baseline, on six of eight MuJoCo tasks. With twin-$Q$ labels, RDA2C matches SAC at matched batch size and update frequency.