The Curse of Memory in Stochastic Approximation

📅 2023-09-06
🏛️ IEEE Conference on Decision and Control
📈 Citations: 9
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper investigates the convergence of constant-step-size stochastic approximation (SA) algorithms under Markovian noise, focusing on root-finding—i.e., solving (f( heta^*) = 0)—and precisely characterizing the inherent bias and covariance error. To overcome the limitations of the classical i.i.d. noise assumption, we propose a joint parameter-perturbation process framework grounded in geometric ergodicity. This enables the first systematic analysis revealing a non-zero steady-state bias induced by “memory effects” in Markov noise, for which we derive a closed-form expression. Concurrently, we establish an explicit (O(alpha)) upper bound on the covariance error, quantifying how Markov dependence amplifies estimation error. Our theoretical results rigorously apply to temporal modeling settings such as TD-learning, and are corroborated by numerical experiments.
📝 Abstract
Theory and application of stochastic approximation (SA) has grown within the control systems community since the earliest days of adaptive control. This paper takes a new look at the topic, motivated by recent results establishing remarkable performance of SA with (sufficiently small) constant step-size $alpha >0$. If averaging is implemented to obtain the final parameter estimate, then the estimates are asymptotically unbiased with nearly optimal asymptotic covariance. These results have been obtained for random linear SA recursions with i.i.d. coefficients. This paper obtains very different conclusions in the more common case of geometrically ergodic Markovian disturbance: (i) The target bias is identified, even in the case of non-linear SA, and is in general non-zero. The remaining results are established for linear SA recursions: (ii) the bivariate parameter-disturbance process is geometrically ergodic in a topological sense; (iii) the representation for bias has a simpler form in this case, and cannot be expected to be zero if there is multiplicative noise; (iv) the asymptotic covariance of the averaged parameters is within $O(alpha)$ of optimal. The error term is identified, and may be massive if mean dynamics are not well conditioned. The theory is illustrated with application to TD-learning.
Problem

Research questions and friction points this paper is trying to address.

Analyzing constant step-size stochastic approximation algorithms for root finding
Studying convergence and bias in optimization and reinforcement learning
Evaluating performance of Polyak-Ruppert averaging with fixed step-size
Innovation

Methods, ideas, or system contributions that make the work stand out.

Constant step-size stochastic approximation for optimization
Geometric ergodicity of Markov chain parameter process
Polyak-Ruppert averaging achieves near-optimal covariance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
University of Florida
C
Caio Kalil Lauand
Department of ECE, University of Florida, Gainesville
S
Sean P. Meyn
Department of ECE, University of Florida, Gainesville