🤖 AI Summary
This paper addresses error propagation in discrete-time stochastic optimal control under the dynamic programming (DP) framework, specifically analyzing how estimation errors accumulate and amplify during backward recursion.
Method: We propose a natural decomposition of value function estimation errors, integrating kernel ridge regression in reproducing kernel Hilbert spaces with Monte Carlo simulation to rigorously characterize both per-step estimation errors and their backward propagation dynamics.
Contribution/Results: We establish the first theoretical framework for error propagation analysis in DP, proving that accumulated error grows controllably with the number of backward steps. The method significantly improves accuracy and numerical stability in American option pricing—demonstrating robust convergence even in high-dimensional and non-Markovian settings. By unifying theoretical rigor with computational feasibility, our approach provides a novel paradigm for financial derivative pricing and policy optimization.
📝 Abstract
This paper investigates theoretical and methodological foundations for stochastic optimal control (SOC) in discrete time. We start formulating the control problem in a general dynamic programming framework, introducing the mathematical structure needed for a detailed convergence analysis. The associate value function is estimated through a sequence of approximations combining nonparametric regression methods and Monte Carlo subsampling. The regression step is performed within reproducing kernel Hilbert spaces (RKHSs), exploiting the classical KRR algorithm, while Monte Carlo sampling methods are introduced to estimate the continuation value. To assess the accuracy of our value function estimator, we propose a natural error decomposition and rigorously control the resulting error terms at each time step. We then analyze how this error propagates backward in time-from maturity to the initial stage-a relatively underexplored aspect of the SOC literature. Finally, we illustrate how our analysis naturally applies to a key financial application: the pricing of American options.