Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文提出正则化强调时序差分学习(RETD)方法,解决ETD在常数步长下稳定性问题,通过引入延迟修正机制改善了采样动态。
📝 Abstract
Emphatic temporal-difference learning (ETD) stabilizes the expected off-policy TD update and changes its projection geometry, but neither property determines constant-stepsize sampled dynamics. We construct an ergodic two-state counterexample in which the ETD mean map contracts while the sampled product has a positive top Lyapunov exponent. Regenerative-cycle analysis separates this sign from the infinite variance of the follow-on trace. We introduce regularized emphatic TD (RETD), a normalized first-order post-shock repair that leaves the trace and importance ratios unchanged, stores the emphatic TD signal in a leaky scalar state, and releases a delayed correction. RETD's raw equilibrium is an affine shift of the ETD equilibrium; single- and two-regularization readouts recover the ETD fixed point exactly. We prove almost-sure convergence for harmonic diminishing stepsizes and a conditional constant-stepsize moment-contraction result from a Markovian random-product bound. RETD has certified negative exponents on the two-state construction and one Baird point, whereas the positive Baird ETD sign remains numerical. Paired 10,000-run experiments validate both separations, fixed-point recovery, a nonmonotone stability region, and task dependence. RETD changes post-shock dynamics; it does not reduce the shared follow-on-trace variance.
Problem

Research questions and friction points this paper is trying to address.

Emphatic Temporal-Difference Learning
Constant Stepsizes
Sampled Dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Regularized Emphatic TD (RETD)
Leaky Scalar State
Negative Lyapunov Exponents
Almost-sure Convergence
🔎 Similar Papers