🤖 AI Summary
Soft Actor-Critic (SAC) for discrete action spaces suffers from weak in-distribution and out-of-distribution generalization, low sample efficiency, and insufficient robustness for safe deployment. Method: We propose a maximum-entropy reinforcement learning framework with statistical constraints, introducing— for the first time in discrete SAC—a proxy critic-guided statistical regularization mechanism that explicitly enforces distributional robustness of the policy entropy objective, thereby mitigating domain shift effects. Contribution/Results: Evaluated on Atari 2600 under low-data regimes, our method achieves an average performance gain of 12.7% over baselines on both in-distribution and out-of-distribution tasks. It significantly improves policy generalization and deployment robustness, establishing a novel paradigm for discrete control under real-world constraints—namely, limited samples and dynamically shifting environments.
📝 Abstract
We present a novel extension to the family of Soft Actor-Critic (SAC) algorithms. We argue that based on the Maximum Entropy Principle, discrete SAC can be further improved via additional statistical constraints derived from a surrogate critic policy. Furthermore, our findings suggests that these constraints provide an added robustness against potential domain shifts, which are essential for safe deployment of reinforcement learning agents in the real-world. We provide theoretical analysis and show empirical results on low data regimes for both in-distribution and out-of-distribution variants of Atari 2600 games.