Score
Designs, implements, and analyzes algorithms and training procedures that jointly estimate policy (actor) and value (critic) components, covering on- and off-policy actor-critic methods, schedule-based training, and continuous-time or SDE-aware variants. Builds update rules, architectures, and training schedules that support entropy-regularized policy improvement, flow-based and distributional policy/return representations (e.g., conditional flow matching, scale-distributional approaches), multimodal or Gaussian policy parameterizations, and implementations for continuous control and return-distribution estimation.
This study addresses the challenges of deploying Actor-Critic algorithms in real-world control systems, where poor reliability and high sensitivity to hyperparameters often hinder practical application. Focusing on a real-world water treatment plant control task, the authors conduct over 33,000 large-scale ablation experiments to systematically evaluate how key algorithmic components—such as policy update schemes, action distribution representations, gradient estimation methods, and update frequencies—affect performance stability and hyperparameter robustness. Their empirical analysis reveals, for the first time, that commonly adopted default configurations (e.g., Gaussian action distributions with pathwise derivatives) exhibit low reliability, whereas bounded action distributions combined with adaptive update strategies substantially enhance robustness. The work identifies high-stability algorithmic configurations that significantly reduce performance variance under limited tuning budgets, offering actionable, component-level design guidelines for industrial deployment.
This work addresses the challenge in offline reinforcement learning that multimodal data distributions cannot be effectively modeled by conventional Gaussian policies. To this end, the authors propose Flow Actor-Critic (FAC), a novel method that, for the first time, integrates normalizing flows into both the policy network and the conservative critic design. Specifically, the approach leverages flow-based models to construct a highly expressive policy capable of accurately capturing complex behavioral distributions. Furthermore, it introduces a flow-based behavioral proxy model to derive a new regularizer for the critic, thereby enhancing the conservatism of value estimation. The proposed method achieves state-of-the-art performance on established offline reinforcement learning benchmarks, including D4RL and OGBench.
To address value overestimation caused by out-of-distribution actions and the difficulty of end-to-end optimization in KL-constrained policy iteration for offline reinforcement learning, this paper reformulates KL-regularized policy updates as a differentiable diffusion-based noise regression task—enabling full diffusion-model parameterization of the target policy. We introduce a soft Q-gradient guidance mechanism, jointly integrated with Q-function ensembling and lower-confidence-bound (LCB) estimation, to preserve policy multimodality while enhancing training stability. Our approach unifies diffusion models, the Actor-Critic framework, and KL-constrained optimization into a single coherent architecture. Evaluated on the D4RL benchmark, it consistently outperforms existing methods across nearly all tasks, achieving state-of-the-art performance.
This work addresses the long-standing open problem of global convergence for single-sample, single-timescale Actor-Critic algorithms in continuous state-action spaces. Using the linear quadratic regulator (LQR) as a canonical model, we establish the first global convergence guarantee to an ε-optimal policy. Our analysis integrates tools from control theory (exploiting the analytic structure of LQR), stochastic approximation, policy gradient estimation, and nonconvex optimization. We rigorously prove that the algorithm converges to an ε-optimal policy with sample complexity O(ε⁻²), matching the information-theoretic lower bound in order. This result breaks prior theoretical dependencies on discrete state-action spaces or two-timescale stepsize regimes. It provides the first tight convergence guarantee for widely deployed single-timescale Actor-Critic methods in continuous domains, thereby bridging a critical gap between theoretical analysis and practical reinforcement learning applications.
This paper addresses off-policy soft-maximum Actor-Critic algorithms under state distribution mismatch, where conventional density-ratio correction is infeasible or impractical. Method: We propose a unified finite-sample analysis framework for stochastic approximation algorithms operating on time-varying Markov chains, employing softmax policy parameterization, single-step stochastic updates, and an inexact critic—without requiring density-ratio correction, stationarity assumptions, or exact gradient access. Contribution/Results: For tabular MDPs, we establish the first global optimality guarantee for such off-policy Actor-Critic methods under these weak conditions. Our novel uniform contraction analysis tool enables rigorous finite-sample convergence characterization, yielding an $O(1/sqrt{T})$ rate. This significantly relaxes classical strong assumptions (e.g., ergodicity, exact gradients, or importance sampling), thereby enhancing theoretical interpretability and practical relevance to deep reinforcement learning training dynamics.
This work investigates the impact of the Critic on policy update variance and convergence in entropy-regularized Actor-Critic algorithms. Under a finite-horizon, discounted setting with entropy regularization, we provide the first rigorous proof that an exact Critic, when used as a baseline, substantially reduces the variance of policy gradients; moreover, even with small approximation errors, it still ensures rapid convergence. Stochastic gradient analysis reveals that, given an exact Critic, the algorithm achieves an ε-optimal regularized value function with only Õ(log(1/ε)) samples, matching the sample complexity of deterministic policy gradient methods. These findings underscore the critical importance of prioritizing accurate Critic learning in such frameworks.
This work addresses the challenges of optimizing dynamic risk measures—such as expectiles and Conditional Value-at-Risk (CVaR)—under stochastic policies, where conventional policy gradient methods require perturbations to state transitions and value estimation typically relies on environment models. To overcome these limitations, the authors propose a novel policy gradient approach that eliminates the need for transition perturbations and leverages elicitable risk statistics to enable model-free value learning for dynamic risk. Building upon Expected SARSA, they develop an off-policy Actor-Critic algorithm that integrates these components within a unified framework. This study presents the first method capable of estimating policy gradients for dynamic risk without requiring transition perturbations, thereby introducing elicitable risk measures into model-free reinforcement learning. Empirical results demonstrate that the proposed approach effectively learns risk-sensitive policies that avoid hazardous outcomes, significantly outperforming existing baselines in relevant tasks.
This work proposes Active Importance Sampling Actor-Critic (AISAC), a novel algorithm addressing the high variance inherent in policy gradient methods, which often leads to inefficient learning and unstable training. AISAC treats the behavior policy as a learnable component and dynamically adapts its distribution via importance sampling to align with the target policy gradient direction, thereby minimizing estimator variance while preserving unbiasedness. Leveraging a Gaussian behavior policy optimized through cross-entropy minimization, the method enables efficient policy updates and value estimation in continuous control tasks. Empirical results demonstrate that AISAC significantly improves learning speed, sample efficiency, and training stability on benchmarks such as Inverted Pendulum and Half Cheetah, while exhibiting strong robustness to hyperparameter variations.
This work addresses the lack of direct metrics and control mechanisms for critic model complexity in conventional actor-critic reinforcement learning. It introduces spectral effective rank entropy—a novel measure derived from singular value decomposition—to quantify the complexity of critic weight matrices. The relationship between this complexity metric and training dynamics is empirically analyzed within TD3 and PPO algorithms. Building on these insights, the study proposes a spectral entropy regularization term to actively regulate critic complexity. Experimental results demonstrate that critic complexity can be reliably measured and effectively controlled, with its impact on performance exhibiting strong task dependence. These findings reveal heterogeneous interactions among model complexity, algorithmic design, task characteristics, and hyperparameter choices.