A Concentration Bound for Two-Timescale Actor-Critic Algorithm

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of uniform anytime bounds in existing two-timescale Actor-Critic algorithms, which hinders the characterization of finite-time behavior under function approximation in the average-reward setting. To overcome this limitation, the work integrates stochastic approximation theory with concentration inequalities to derive unified concentration bounds for the algorithm under function approximation, while analyzing the high-probability evolution trajectory of the Actor parameters. The primary contribution is establishing the first uniform anytime bound for this class of algorithms, demonstrating that the parameters enter and remain within a safe region with high probability. Furthermore, an explicit upper bound on the Actor error is provided, and empirical experiments confirm its monotonic decrease over iterations.
📝 Abstract
Significant research effort has been directed in recent years towards establishing both asymptotic and non-asymptotic convergence guarantees for two-timescale actor--critic algorithms, where the actor recursion is run on a slower timescale than the critic recursion. This work derives a uniform all-time concentration bound for the actor--critic algorithm with function approximation in the long-run average-reward setting. This bound helps us analyze the behavior of the actor parameter with high probability. We show that, after some finite time, the actor parameter enters a safe region and remains within it thereafter with high probability. Specifically, with probability at least $1-ε_1-ε_2$, the actor error $\Vert θ_k-θ^{*}\Vert$ is $O\left(\frac{n_0^{3/4}}{k}\frac{1}{\sqrt{ε_2}}+\left(\frac{1}{n_0}\right)^{1/4}\log^{1/4}\left(\frac{1}{ε_1}\right)+\left(\frac{1}{n_0}\right)^{1/4}\right)$ for all $k\geq n_0$ and sufficiently large $n_0$. We also present experimental results demonstrating that the aforementioned actor error diminishes with the number of actor-parameter updates.
Problem

Research questions and friction points this paper is trying to address.

Two-Timescale Actor-Critic
Concentration Bound
Function Approximation
Average-Reward
Actor Error
Innovation

Methods, ideas, or system contributions that make the work stand out.

Two-Timescale Actor-Critic
Concentration Bound
Function Approximation
Average-Reward
Non-asymptotic Convergence
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.