q-value estimation

Designs and builds estimators of the action-value (Q) function from observed transition data, including fitted-Q procedures for policy evaluation and offline Q estimation. Also develops optimism-based Q estimators and data-dependent transition bonuses or correction terms to control bias, regularize exploration, and enable analyses of estimation error and data-dependent regret tied to loss-prediction or function-approximation error.

q-valueestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$201K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Statistical inference of the value function for reinforcement learning in infinite‐horizon settings

Jan 13, 2020
CS
C. Shi
🏛️ London School of Economics and Political Science | North Carolina State University

In infinite-horizon reinforcement learning, statistical inference for policy value functions remains challenging—particularly under high-dimensional state spaces, non-unique optimal policies, and infinitely many decision points—making reliable confidence interval construction difficult. To address this, we propose the first series-expansion-based Q-function modeling framework and the recursive Series-based Adaptive Value Estimation (SAVE) method. SAVE integrates sieve estimation with asymptotic statistical inference to ensure nominal coverage probability under policy-dependent data. We establish theoretical guarantees showing that the constructed confidence intervals achieve asymptotically exact coverage. Simulation studies demonstrate robustness to model misspecification and policy variation. Empirical evaluation on a mobile health dataset confirms the statistically significant improvement in patient health outcomes attributable to the RL intervention.

Construct confidence intervals for policy value in infinite horizon settingsDevelop SAVE method for recursive policy and value updates with dataEstimate action-value function using series/sieve methods for accurate CIs

Orthogonalized Estimation of Difference of $Q$-functions

Jun 12, 2024
DC
Defu Cao
🏛️ University of Southern California

In offline reinforcement learning, accurately estimating the action-value function difference $Q^pi(s,1) - Q^pi(s,0)$ from static historical data—enabling optimal multi-action policy selection—is a fundamental challenge. This paper introduces the first dynamic extension of the R-learner from causal inference to RL, proposing an orthogonalized estimation framework for offline Q-difference estimation. The method is robust to slowly converging auxiliary models (e.g., black-box Q-function and behavior policy estimators), ensures consistent policy optimization under mild marginal conditions, and achieves faster convergence rates. Theoretically, it establishes estimation consistency without requiring strong parametric assumptions on auxiliary models. Empirically, the approach directly enables optimal multi-action policy selection and significantly improves offline decision-making performance across benchmark tasks.

Estimates difference of Q-functions for offline RLImproves convergence with nuisance estimationOptimizes policies using orthogonal learning

Post Reinforcement Learning Inference

Feb 17, 2023
VS
Vasilis Syrgkanis
🏛️ Stanford University | Hong Kong University of Science and Technology

In reinforcement learning, adaptive interaction data—where the behavior policy is nonstationary—invalidates standard estimators, undermining asymptotic normality for off-policy counterfactual policy evaluation and dynamic treatment effect (DTE) inference. To address this, we propose a weighted Z-estimation framework that constructs time-varying adaptive weights to stabilize heteroskedasticity, achieving, for the first time in the RL off-policy setting, both consistent and asymptotically normal DTE estimation. Our approach integrates dynamic causal inference with asymptotic statistical theory, enabling rigorous hypothesis testing and construction of uniformly valid confidence regions. Simulation studies and real-world RL experiments demonstrate substantial improvements in confidence interval coverage and statistical power. The method provides the first solution for structural parameter inference under adaptive experimentation that simultaneously offers theoretical guarantees—namely consistency, asymptotic normality, and uniform validity—and empirical robustness.

Address nonstationary variance in adaptive reinforcement learning environmentsDevelop weighted Z-estimation for dynamic treatment effect analysisEstimate counterfactual policies post reinforcement learning data collection

Robust Fitted-Q-Evaluation and Iteration under Sequentially Exogenous Unobserved Confounders

Feb 01, 2023
DB
David Bruns-Smith
🏛️ University of California Berkeley | University of Southern California

Offline reinforcement learning (RL) in high-stakes domains such as medicine suffers from sequential exogenous unobserved confounding, inducing bias in policy evaluation and optimization. This work addresses robust policy evaluation and optimization under a sensitivity model. We propose Orthogonalized Robust Fitted Q-Iteration (ORFQI), the first algorithm to embed the closed-form solution of the robust Bellman operator into a loss minimization framework via orthogonalization. Additionally, we introduce bias-corrected quantile regression to enhance statistical robustness and computational efficiency. Theoretically, we derive finite-sample complexity bounds for the estimator. Empirically, we validate our approach on real-world longitudinal sepsis data, demonstrating significantly improved robustness in policy evaluation. Moreover, our conservative confidence bounds enable warm-starting optimistic online RL, facilitating safer deployment in clinical settings.

Addresses offline RL with unobserved confounders in sequential decision-makingEnables valid policy optimization from healthcare and economic observational datasetsProposes robust policy evaluation under sensitivity models for observational data

Asymptotic Theory for IV-Based Reinforcement Learning with Potential Endogeneity

Mar 06, 2021
JL
Jin Li
🏛️ The University of Hong Kong | The Hong Kong University of Science and Technology

This paper addresses the reinforcement bias and endogeneity arising from the dynamic coupling between data generation and policy evaluation in reinforcement learning (RL). We propose Instrumental Variable RL (IV-RL), the first RL paradigm explicitly designed to handle endogeneity via instrumental variables. By modeling policy iteration as a Markov process with instrument-dependent transitions, we develop an asymptotic statistical theory for IV-RL, derive an inferentially valid estimator of the optimal policy, and quantify how temporal dependence degrades inference accuracy. Our method integrates instrumental variable estimation, stochastic approximation theory, and Markov decision process (MDP) modeling. We rigorously establish strong consistency and asymptotic normality of the IV-RL estimator, enabling unbiased policy evaluation and principled confidence interval construction. The key contribution is breaking the conventional offline RL assumption of exogenous data: IV-RL provides the first statistically grounded, endogeneity-robust correction framework for dynamic decision-making under endogenous feedback.

Data AnalysisMarkov Decision ProcessReinforcement Bias

Latest Papers

What's happening recently
View more

This work addresses the severe overestimation bias in Q-learning within large discrete action spaces, which stems from high variance in Q-value estimation. Existing approaches—either fully coupled or fully decoupled—are respectively hindered by persistent positive or negative biases. To overcome this limitation, the paper introduces an action intersection strategy that dynamically controls the proportion of trajectory data shared between two Q-functions: updates are coupled on shared samples and decoupled on non-shared ones, thereby establishing a semi-decoupled mechanism for fine-grained control over both overestimation and underestimation biases. This approach breaks free from conventional paradigms by enabling flexible bias-direction balancing in large action spaces for the first time. Empirical results demonstrate significant performance gains over current state-of-the-art methods in both tabular and deep reinforcement learning settings, while also uncovering the underlying mechanisms driving this improvement.

bias bottlenecklarge discrete action spaceoverestimation bias

This study addresses the lack of finite-sample theory and the challenge of greedy optimization in functional action spaces for offline reinforcement learning by investigating the theoretical foundations of functional-action Fitted Q-Iteration (FQI). Methodologically, it introduces a critic relative coverage condition to replace conventional density assumptions, thereby overcoming bottlenecks in convergence analysis. Efficient Q-value estimation is achieved by integrating smoothed regularized policy search, kernel ridge regression, and adaptive neural networks. Theoretically, the work establishes polynomially decaying regret bounds. Empirically, experiments demonstrate that functional policies outperform constant baselines and validate the effectiveness of the proposed coverage condition. Overall, this research provides both theoretical and algorithmic support for offline decision-making in functional action spaces.

Coverage conditionFinite-sample theoryFitted Q-iteration

This work addresses the challenge of Q-value overestimation in offline reinforcement learning, which arises from distributional shift between the behavior policy and the learning policy and can mislead policy optimization. To mitigate this issue, the paper proposes Behavior-Advantage Corrected Policy Evaluation (BAC-PE), a novel approach that leverages the Q-function of the behavior policy to correct the Q-estimates of the learning policy. BAC-PE further integrates a diffusion model for policy representation, enhancing expressiveness and training stability through Q-guided sampling and distribution-matching regularization. Theoretical analysis establishes the convergence of BAC-PE and provides an upper bound on its estimation error. Empirical evaluation on the D4RL benchmark demonstrates that the proposed method significantly outperforms existing offline reinforcement learning algorithms, achieving state-of-the-art performance.

behavioral policydistribution shiftoffline reinforcement learning

Hot Scholars

MM

Mahmoud M. Salim

Electronics and Electrical Communication, October 6 University
6GResource AllocationOptimizationML
CS

Chengchun Shi

London School of Economics and Political Science
Large Language ModelsReinforcement LearningStatistics
NK

Nathan Kallus

Cornell University
Optimization under uncertaintyCausal inferenceBanditsRL
ZW

Zhenke Wu

Associate Professor of Biostatistics (with tenure), University of Michigan
StatisticsCausalityDigital HealthPrecision Health
LV

Lars van der Laan

Ph.D. Student in Statistics, University of Washington, Seattle
StatisticsMachine LearningCausal Inference