Score
Designs and builds estimators of the action-value (Q) function from observed transition data, including fitted-Q procedures for policy evaluation and offline Q estimation. Also develops optimism-based Q estimators and data-dependent transition bonuses or correction terms to control bias, regularize exploration, and enable analyses of estimation error and data-dependent regret tied to loss-prediction or function-approximation error.
In infinite-horizon reinforcement learning, statistical inference for policy value functions remains challenging—particularly under high-dimensional state spaces, non-unique optimal policies, and infinitely many decision points—making reliable confidence interval construction difficult. To address this, we propose the first series-expansion-based Q-function modeling framework and the recursive Series-based Adaptive Value Estimation (SAVE) method. SAVE integrates sieve estimation with asymptotic statistical inference to ensure nominal coverage probability under policy-dependent data. We establish theoretical guarantees showing that the constructed confidence intervals achieve asymptotically exact coverage. Simulation studies demonstrate robustness to model misspecification and policy variation. Empirical evaluation on a mobile health dataset confirms the statistically significant improvement in patient health outcomes attributable to the RL intervention.
In offline reinforcement learning, accurately estimating the action-value function difference $Q^pi(s,1) - Q^pi(s,0)$ from static historical data—enabling optimal multi-action policy selection—is a fundamental challenge. This paper introduces the first dynamic extension of the R-learner from causal inference to RL, proposing an orthogonalized estimation framework for offline Q-difference estimation. The method is robust to slowly converging auxiliary models (e.g., black-box Q-function and behavior policy estimators), ensures consistent policy optimization under mild marginal conditions, and achieves faster convergence rates. Theoretically, it establishes estimation consistency without requiring strong parametric assumptions on auxiliary models. Empirically, the approach directly enables optimal multi-action policy selection and significantly improves offline decision-making performance across benchmark tasks.
In reinforcement learning, adaptive interaction data—where the behavior policy is nonstationary—invalidates standard estimators, undermining asymptotic normality for off-policy counterfactual policy evaluation and dynamic treatment effect (DTE) inference. To address this, we propose a weighted Z-estimation framework that constructs time-varying adaptive weights to stabilize heteroskedasticity, achieving, for the first time in the RL off-policy setting, both consistent and asymptotically normal DTE estimation. Our approach integrates dynamic causal inference with asymptotic statistical theory, enabling rigorous hypothesis testing and construction of uniformly valid confidence regions. Simulation studies and real-world RL experiments demonstrate substantial improvements in confidence interval coverage and statistical power. The method provides the first solution for structural parameter inference under adaptive experimentation that simultaneously offers theoretical guarantees—namely consistency, asymptotic normality, and uniform validity—and empirical robustness.
Offline reinforcement learning (RL) in high-stakes domains such as medicine suffers from sequential exogenous unobserved confounding, inducing bias in policy evaluation and optimization. This work addresses robust policy evaluation and optimization under a sensitivity model. We propose Orthogonalized Robust Fitted Q-Iteration (ORFQI), the first algorithm to embed the closed-form solution of the robust Bellman operator into a loss minimization framework via orthogonalization. Additionally, we introduce bias-corrected quantile regression to enhance statistical robustness and computational efficiency. Theoretically, we derive finite-sample complexity bounds for the estimator. Empirically, we validate our approach on real-world longitudinal sepsis data, demonstrating significantly improved robustness in policy evaluation. Moreover, our conservative confidence bounds enable warm-starting optimistic online RL, facilitating safer deployment in clinical settings.
This paper addresses the reinforcement bias and endogeneity arising from the dynamic coupling between data generation and policy evaluation in reinforcement learning (RL). We propose Instrumental Variable RL (IV-RL), the first RL paradigm explicitly designed to handle endogeneity via instrumental variables. By modeling policy iteration as a Markov process with instrument-dependent transitions, we develop an asymptotic statistical theory for IV-RL, derive an inferentially valid estimator of the optimal policy, and quantify how temporal dependence degrades inference accuracy. Our method integrates instrumental variable estimation, stochastic approximation theory, and Markov decision process (MDP) modeling. We rigorously establish strong consistency and asymptotic normality of the IV-RL estimator, enabling unbiased policy evaluation and principled confidence interval construction. The key contribution is breaking the conventional offline RL assumption of exogenous data: IV-RL provides the first statistically grounded, endogeneity-robust correction framework for dynamic decision-making under endogenous feedback.
This work addresses the severe overestimation bias in Q-learning within large discrete action spaces, which stems from high variance in Q-value estimation. Existing approaches—either fully coupled or fully decoupled—are respectively hindered by persistent positive or negative biases. To overcome this limitation, the paper introduces an action intersection strategy that dynamically controls the proportion of trajectory data shared between two Q-functions: updates are coupled on shared samples and decoupled on non-shared ones, thereby establishing a semi-decoupled mechanism for fine-grained control over both overestimation and underestimation biases. This approach breaks free from conventional paradigms by enabling flexible bias-direction balancing in large action spaces for the first time. Empirical results demonstrate significant performance gains over current state-of-the-art methods in both tabular and deep reinforcement learning settings, while also uncovering the underlying mechanisms driving this improvement.
This study addresses the lack of finite-sample theory and the challenge of greedy optimization in functional action spaces for offline reinforcement learning by investigating the theoretical foundations of functional-action Fitted Q-Iteration (FQI). Methodologically, it introduces a critic relative coverage condition to replace conventional density assumptions, thereby overcoming bottlenecks in convergence analysis. Efficient Q-value estimation is achieved by integrating smoothed regularized policy search, kernel ridge regression, and adaptive neural networks. Theoretically, the work establishes polynomially decaying regret bounds. Empirically, experiments demonstrate that functional policies outperform constant baselines and validate the effectiveness of the proposed coverage condition. Overall, this research provides both theoretical and algorithmic support for offline decision-making in functional action spaces.
研究通过近似最大贝尔曼算子并使用Neyman正交性提出去偏估计器,解决强化学习中的最优值离线推理问题。
This work addresses the challenge of Q-value overestimation in offline reinforcement learning, which arises from distributional shift between the behavior policy and the learning policy and can mislead policy optimization. To mitigate this issue, the paper proposes Behavior-Advantage Corrected Policy Evaluation (BAC-PE), a novel approach that leverages the Q-function of the behavior policy to correct the Q-estimates of the learning policy. BAC-PE further integrates a diffusion model for policy representation, enhancing expressiveness and training stability through Q-guided sampling and distribution-matching regularization. Theoretical analysis establishes the convergence of BAC-PE and provides an upper bound on its estimation error. Empirical evaluation on the D4RL benchmark demonstrates that the proposed method significantly outperforms existing offline reinforcement learning algorithms, achieving state-of-the-art performance.
本文针对强化学习中策略评估的高方差问题,提出了一种双循环梯度算法来学习对转移不确定性鲁棒的行为策略。