π€ AI Summary
This work establishes the first global concentration boundβvalid for all $ n geq n_0 $βon the finite-sample behavior of online single-trajectory TD(0) with linear function approximation under Markovian noise. Unlike offline settings relying on i.i.d. samples or ergodic state visits, it directly addresses non-independent, non-stationary, temporally dependent data. Methodologically, it integrates stochastic approximation theory with Poisson equation analysis to characterize Markovian bias and employs a relaxed concentration inequality to circumvent the lack of a priori boundedness in iterates. The result delivers an explicit exponential decay probability bound, rigorously quantifying the convergence rate of estimation error with respect to iteration count. This provides the first finite-sample theoretical guarantee for online reinforcement learning with fully explicit constant dependencies.
π Abstract
We derive a concentration bound of the type `for all $n geq n_0$ for some $n_0$' for TD(0) with linear function approximation. We work with online TD learning with samples from a single sample path of the underlying Markov chain. This makes our analysis significantly different from offline TD learning or TD learning with access to independent samples from the stationary distribution of the Markov chain. We treat TD(0) as a contractive stochastic approximation algorithm, with both martingale and Markov noises. Markov noise is handled using the Poisson equation and the lack of almost sure guarantees on boundedness of iterates is handled using the concept of relaxed concentration inequalities.