🤖 AI Summary
This study addresses the high mean squared error and inherent intractability of empirically estimating the invariant distribution of Markov chains. To this end, it proposes lookahead and plug-in lookahead estimators that leverage kernel propagation to optimize estimation accuracy. Theoretically, we prove that multi-step lookahead is asymptotically superior to empirical estimation under reversible kernels, and establish necessary and sufficient conditions under which the plug-in estimator outperforms its empirical counterpart. Empirically, the proposed methods are validated through Monte Carlo simulations and real-world mobility data from the Lyon metropolitan area. The results demonstrate that our approach significantly reduces mean squared error and enhances estimation efficiency, offering a novel paradigm for the high-precision estimation of Markov chain invariant distributions.
📝 Abstract
We consider the estimation of the invariant distribution $\pi$ of a Markov kernel on a discrete state space (or of the mean value $\pi(f)$ of a functional $f$ under $\pi$) when solving the invariance equation is not feasible. Look-ahead estimators exploit the knowledge of the kernel by propagating the empirical occupation measure through $k$ steps of the chain. When the kernel is known, as in Markov chain Monte Carlo, we compare their mean squared error with that of the empirical estimator, for independent or Markovian data, stationary or not. Our main result establishes that, for Markovian data, the one-step estimator always improves on the empirical one asymptotically, and that the improvement grows with $k$ for reversible kernels. A non-reversible counterexample shows that reversibility cannot simply be dropped. In the statistical setting where the kernel belongs to a parametric family with unknown parameter, we introduce plug-in look-ahead estimators, which perform the look-ahead step with an estimated kernel. When the parameter is estimated at the parametric rate, we give a necessary and sufficient condition under which the plug-in look-ahead estimator outperforms the empirical one. The theoretical results are illustrated by extensive simulation studies, which also shed further light on the behaviour of the estimators at hand. Both settings are finally applied to real data arising from geographical mobility problems in the Lyon metropolitan area, France.