Global Optimality and Finite Sample Analysis of Softmax Off-Policy Actor Critic under State Distribution Mismatch

📅 2021-11-04
🏛️ Journal of machine learning research
📈 Citations: 9
✨ Influential: 1
📄 PDF
🤖 AI Summary
This paper addresses off-policy soft-maximum Actor-Critic algorithms under state distribution mismatch, where conventional density-ratio correction is infeasible or impractical. Method: We propose a unified finite-sample analysis framework for stochastic approximation algorithms operating on time-varying Markov chains, employing softmax policy parameterization, single-step stochastic updates, and an inexact critic—without requiring density-ratio correction, stationarity assumptions, or exact gradient access. Contribution/Results: For tabular MDPs, we establish the first global optimality guarantee for such off-policy Actor-Critic methods under these weak conditions. Our novel uniform contraction analysis tool enables rigorous finite-sample convergence characterization, yielding an $O(1/sqrt{T})$ rate. This significantly relaxes classical strong assumptions (e.g., ergodicity, exact gradients, or importance sampling), thereby enhancing theoretical interpretability and practical relevance to deep reinforcement learning training dynamics.
📝 Abstract
In this paper, we establish the global optimality and convergence rate of an off-policy actor critic algorithm in the tabular setting without using density ratio to correct the discrepancy between the state distribution of the behavior policy and that of the target policy. Our work goes beyond existing works on the optimality of policy gradient methods in that existing works use the exact policy gradient for updating the policy parameters while we use an approximate and stochastic update step. Our update step is not a gradient update because we do not use a density ratio to correct the state distribution, which aligns well with what practitioners do. Our update is approximate because we use a learned critic instead of the true value function. Our update is stochastic because at each step the update is done for only the current state action pair. Moreover, we remove several restrictive assumptions from existing works in our analysis. Central to our work is the finite sample analysis of a generic stochastic approximation algorithm with time-inhomogeneous update operators on time-inhomogeneous Markov chains, based on its uniform contraction properties.
Problem

Research questions and friction points this paper is trying to address.

Analyze off-policy actor critic algorithm global optimality
Remove density ratio for state distribution correction
Conduct finite sample analysis with stochastic updates
Innovation

Methods, ideas, or system contributions that make the work stand out.

Off-policy actor critic
Stochastic update step
Learned critic usage
🔎 Similar Papers
2024-05-23Trans. Mach. Learn. Res.Citations: 0
💼 Related Jobs
No related jobs found.
University of Virginia | Microsoft Research Montreal
Shangtong Zhang
Shangtong Zhang
University of Virginia
reinforcement learningstochastic approximation
R
Rémi Tachet des Combes
Microsoft Research Montreal
R
R. Laroche
Microsoft Research Montreal