train actor-critic agents

Designs, implements, and analyzes algorithms and training procedures that jointly estimate policy (actor) and value (critic) components, covering on- and off-policy actor-critic methods, schedule-based training, and continuous-time or SDE-aware variants. Builds update rules, architectures, and training schedules that support entropy-regularized policy improvement, flow-based and distributional policy/return representations (e.g., conditional flow matching, scale-distributional approaches), multimodal or Gaussian policy parameterizations, and implementations for continuous control and return-distribution estimation.

trainactor-criticagents

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$228K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenges of deploying Actor-Critic algorithms in real-world control systems, where poor reliability and high sensitivity to hyperparameters often hinder practical application. Focusing on a real-world water treatment plant control task, the authors conduct over 33,000 large-scale ablation experiments to systematically evaluate how key algorithmic components—such as policy update schemes, action distribution representations, gradient estimation methods, and update frequencies—affect performance stability and hyperparameter robustness. Their empirical analysis reveals, for the first time, that commonly adopted default configurations (e.g., Gaussian action distributions with pathwise derivatives) exhibit low reliability, whereas bounded action distributions combined with adaptive update strategies substantially enhance robustness. The work identifies high-stability algorithmic configurations that significantly reduce performance variance under limited tuning budgets, offering actionable, component-level design guidelines for industrial deployment.

Actor-Criticalgorithm reliabilityhyperparameter sensitivity

This work addresses the challenge in offline reinforcement learning that multimodal data distributions cannot be effectively modeled by conventional Gaussian policies. To this end, the authors propose Flow Actor-Critic (FAC), a novel method that, for the first time, integrates normalizing flows into both the policy network and the conservative critic design. Specifically, the approach leverages flow-based models to construct a highly expressive policy capable of accurately capturing complex behavioral distributions. Furthermore, it introduces a flow-based behavioral proxy model to derive a new regularizer for the critic, thereby enhancing the conservatism of value estimation. The proposed method achieves state-of-the-art performance on established offline reinforcement learning benchmarks, including D4RL and OGBench.

dataset distributionexpressive policiesmulti-modal distributions

To address value overestimation caused by out-of-distribution actions and the difficulty of end-to-end optimization in KL-constrained policy iteration for offline reinforcement learning, this paper reformulates KL-regularized policy updates as a differentiable diffusion-based noise regression task—enabling full diffusion-model parameterization of the target policy. We introduce a soft Q-gradient guidance mechanism, jointly integrated with Q-function ensembling and lower-confidence-bound (LCB) estimation, to preserve policy multimodality while enhancing training stability. Our approach unifies diffusion models, the Actor-Critic framework, and KL-constrained optimization into a single coherent architecture. Evaluated on the D4RL benchmark, it consistently outperforms existing methods across nearly all tasks, achieving state-of-the-art performance.

Enhances policy performance and stability with diffusion models.Formulates KL constraint policy iteration as diffusion noise.Manages out-of-distribution actions in offline RL.

Global Optimality of Single-Timescale Actor-Critic under Continuous State-Action Space: A Study on Linear Quadratic Regulator

Aug 01, 2024
XC
Xuyang Chen
🏛️ National University of Singapore | University of Science and Technology Beijing

This work addresses the long-standing open problem of global convergence for single-sample, single-timescale Actor-Critic algorithms in continuous state-action spaces. Using the linear quadratic regulator (LQR) as a canonical model, we establish the first global convergence guarantee to an ε-optimal policy. Our analysis integrates tools from control theory (exploiting the analytic structure of LQR), stochastic approximation, policy gradient estimation, and nonconvex optimization. We rigorously prove that the algorithm converges to an ε-optimal policy with sample complexity O(ε⁻²), matching the information-theoretic lower bound in order. This result breaks prior theoretical dependencies on discrete state-action spaces or two-timescale stepsize regimes. It provides the first tight convergence guarantee for widely deployed single-timescale Actor-Critic methods in continuous domains, thereby bridging a critical gap between theoretical analysis and practical reinforcement learning applications.

Analyzes single-timescale actor-critic in continuous state-action spaceBridges theory-practice gap in actor-critic performance understandingProves epsilon-optimal solution for linear quadratic regulator (LQR)

Global Optimality and Finite Sample Analysis of Softmax Off-Policy Actor Critic under State Distribution Mismatch

Nov 04, 2021
SZ
Shangtong Zhang
🏛️ University of Virginia | Microsoft Research Montreal

This paper addresses off-policy soft-maximum Actor-Critic algorithms under state distribution mismatch, where conventional density-ratio correction is infeasible or impractical. Method: We propose a unified finite-sample analysis framework for stochastic approximation algorithms operating on time-varying Markov chains, employing softmax policy parameterization, single-step stochastic updates, and an inexact critic—without requiring density-ratio correction, stationarity assumptions, or exact gradient access. Contribution/Results: For tabular MDPs, we establish the first global optimality guarantee for such off-policy Actor-Critic methods under these weak conditions. Our novel uniform contraction analysis tool enables rigorous finite-sample convergence characterization, yielding an $O(1/sqrt{T})$ rate. This significantly relaxes classical strong assumptions (e.g., ergodicity, exact gradients, or importance sampling), thereby enhancing theoretical interpretability and practical relevance to deep reinforcement learning training dynamics.

Analyze off-policy actor critic algorithm global optimalityConduct finite sample analysis with stochastic updatesRemove density ratio for state distribution correction

Latest Papers

What's happening recently
View more

This work investigates the impact of the Critic on policy update variance and convergence in entropy-regularized Actor-Critic algorithms. Under a finite-horizon, discounted setting with entropy regularization, we provide the first rigorous proof that an exact Critic, when used as a baseline, substantially reduces the variance of policy gradients; moreover, even with small approximation errors, it still ensures rapid convergence. Stochastic gradient analysis reveals that, given an exact Critic, the algorithm achieves an ε-optimal regularized value function with only Õ(log(1/ε)) samples, matching the sample complexity of deterministic policy gradient methods. These findings underscore the critical importance of prioritizing accurate Critic learning in such frameworks.

actor-criticcritic estimationentropy regularization

This work addresses the challenges of optimizing dynamic risk measures—such as expectiles and Conditional Value-at-Risk (CVaR)—under stochastic policies, where conventional policy gradient methods require perturbations to state transitions and value estimation typically relies on environment models. To overcome these limitations, the authors propose a novel policy gradient approach that eliminates the need for transition perturbations and leverages elicitable risk statistics to enable model-free value learning for dynamic risk. Building upon Expected SARSA, they develop an off-policy Actor-Critic algorithm that integrates these components within a unified framework. This study presents the first method capable of estimating policy gradients for dynamic risk without requiring transition perturbations, thereby introducing elicitable risk measures into model-free reinforcement learning. Empirical results demonstrate that the proposed approach effectively learns risk-sensitive policies that avoid hazardous outcomes, significantly outperforming existing baselines in relevant tasks.

dynamic riskpolicy gradientrisk-averse

This work proposes Active Importance Sampling Actor-Critic (AISAC), a novel algorithm addressing the high variance inherent in policy gradient methods, which often leads to inefficient learning and unstable training. AISAC treats the behavior policy as a learnable component and dynamically adapts its distribution via importance sampling to align with the target policy gradient direction, thereby minimizing estimator variance while preserving unbiasedness. Leveraging a Gaussian behavior policy optimized through cross-entropy minimization, the method enables efficient policy updates and value estimation in continuous control tasks. Empirical results demonstrate that AISAC significantly improves learning speed, sample efficiency, and training stability on benchmarks such as Inverted Pendulum and Half Cheetah, while exhibiting strong robustness to hyperparameter variations.

Actor-Criticimportance samplingpolicy gradient variance

This work addresses the lack of direct metrics and control mechanisms for critic model complexity in conventional actor-critic reinforcement learning. It introduces spectral effective rank entropy—a novel measure derived from singular value decomposition—to quantify the complexity of critic weight matrices. The relationship between this complexity metric and training dynamics is empirically analyzed within TD3 and PPO algorithms. Building on these insights, the study proposes a spectral entropy regularization term to actively regulate critic complexity. Experimental results demonstrate that critic complexity can be reliably measured and effectively controlled, with its impact on performance exhibiting strong task dependence. These findings reveal heterogeneous interactions among model complexity, algorithmic design, task characteristics, and hyperparameter choices.

actor-critic reinforcement learningcomplexity controlcritic complexity

Hot Scholars

VA

Vaneet Aggarwal

Professor and University Faculty Scholar, Purdue University
Machine LearningReinforcement LearningQuantum ComputingNetworking
ZC

Zhiyong Chen

Shanghai Jiao Tong University
6G networksWireless CommunicationsComputing and Caching Networks
SG

Swetha Ganesh

Purdue University
Reinforcement LearningStochastic OptimizationSpectral Graph Theory
SB

Shalabh Bhatnagar

Professor in the Department of Computer Science and Automation, Indian Institute of Science
Stochastic systemscontrolsimulationoptimization
JW

Jinbo Wen

M.S. Student, Nanjing University of Aeronautics and Astronautics
GenAI+NetworkingContract TheoryMetaverseBlockchain