Score
Design, parameterize, and evaluate stochastic policies that map states or covariates to probability distributions — including function-valued or distributional interventions — by building implementable parametric families or basis representations for policy shifts and constructing estimands and estimators that may avoid reliance on strict positivity. In multi‑agent or adaptive settings, formulate and analyze game‑theoretic reinforcement‑learning and control models to study incentive alignment, strategic equilibria, and the stability of interacting stochastic policies.
Real-world reinforcement learning faces significant challenges due to data scarcity and dynamically changing environments, leading to a growing gap between theoretical advances and practical deployment. This work proposes a practice-oriented, three-stage framework—comprising in-deployment online learning, inter-deployment offline analysis, and multi-round continual optimization—that systematically integrates recent advances in statistical reinforcement learning to enhance data utility, sample efficiency, and deployment strategies. By emphasizing the pivotal role of statistical methods in bridging the theory–practice divide, the framework offers both methodological guidance and novel research directions for developing reinforcement learning systems tailored to real-world scenarios.
This paper addresses the vulnerability of conventional utilitarian policy learning—based on the conditional average treatment effect (CATE)—to outliers and its inability to flexibly accommodate policy caution or leniency under heterogeneous individual treatment effects. We propose an optimal intervention allocation framework grounded in the conditional quantile treatment effect (QoTE). To our knowledge, this is the first work to incorporate distributionally robust welfare into policy learning, formulating a minimax strategy based on QoTE that accommodates the fundamental challenge of non-point-identification of the counterfactual joint distribution. Our method integrates causal inference, distributionally robust optimization, and decision theory, supporting both stochastic and deterministic policies under various identification assumptions. We establish an asymptotically tight upper bound on the regret and demonstrate robustness to model misspecification. The framework generalizes to any welfare objective defined as a functional of the potential outcomes’ joint distribution.
This paper addresses distributionally robust stochastic control in continuous state spaces to mitigate the policy fragility of conventional i.i.d. Markov models—arising from neglecting environmental input distribution shifts and endogenous dependencies. We propose a novel paradigm that balances modeling simplicity and robustness: adaptive adversarial perturbations are embedded within dynamic programming to unify *f*-divergence and Wasserstein-type ambiguity sets. We establish, for the first time in continuous spaces, a unified learning theory for robust value functions under both ambiguity-set classes, providing finite-sample minimax convergence rate bounds. Integrating distributionally robust optimization, stochastic control, and nonparametric statistics, we design a computationally tractable minimax policy learning algorithm. Experiments demonstrate that the framework achieves both statistical efficiency and strong robustness across real-world applications—including supply chain management and finance.
This paper addresses the reinforcement bias and endogeneity arising from the dynamic coupling between data generation and policy evaluation in reinforcement learning (RL). We propose Instrumental Variable RL (IV-RL), the first RL paradigm explicitly designed to handle endogeneity via instrumental variables. By modeling policy iteration as a Markov process with instrument-dependent transitions, we develop an asymptotic statistical theory for IV-RL, derive an inferentially valid estimator of the optimal policy, and quantify how temporal dependence degrades inference accuracy. Our method integrates instrumental variable estimation, stochastic approximation theory, and Markov decision process (MDP) modeling. We rigorously establish strong consistency and asymptotic normality of the IV-RL estimator, enabling unbiased policy evaluation and principled confidence interval construction. The key contribution is breaking the conventional offline RL assumption of exogenous data: IV-RL provides the first statistically grounded, endogeneity-robust correction framework for dynamic decision-making under endogenous feedback.
This work addresses stochastic team games under unknown dynamics. Methodologically, it proposes the Logit-Q dynamics framework—the first to couple Logit-response dynamics with Q-learning within an auxiliary stage game—where Q-functions drive state-dependent payoffs to enable efficient equilibrium learning. Technically, it introduces a novel analytical approach combining fictitious static Q-estimation scenarios with asymptotic coupling to the true dynamic environment, integrated with slowly varying epoch scheduling and coupling-based convergence analysis. This yields the first convergence and rationality guarantees for non-fully controllable stochastic games. Theoretically, the algorithm converges to an approximately optimal team equilibrium, with quantifiable approximation error; exhibits rationality against pure stationary-strategy opponents; and retains convergence when stage payoffs form a potential game and state transitions are controlled by a single agent.
In reinforcement learning, adaptive interaction data—where the behavior policy is nonstationary—invalidates standard estimators, undermining asymptotic normality for off-policy counterfactual policy evaluation and dynamic treatment effect (DTE) inference. To address this, we propose a weighted Z-estimation framework that constructs time-varying adaptive weights to stabilize heteroskedasticity, achieving, for the first time in the RL off-policy setting, both consistent and asymptotically normal DTE estimation. Our approach integrates dynamic causal inference with asymptotic statistical theory, enabling rigorous hypothesis testing and construction of uniformly valid confidence regions. Simulation studies and real-world RL experiments demonstrate substantial improvements in confidence interval coverage and statistical power. The method provides the first solution for structural parameter inference under adaptive experimentation that simultaneously offers theoretical guarantees—namely consistency, asymptotic normality, and uniform validity—and empirical robustness.
This work addresses the lack of a unified theoretical framework for reinforcement learning, which has hindered systematic analysis of its convergence, sample complexity, and generalization. Building upon Markov decision processes and Bellman operators, the paper introduces a cohesive analytical framework that integrates tools from operator theory, stochastic approximation, convex duality, and function approximation. This framework encompasses a broad range of algorithms, including value iteration, policy iteration, temporal difference methods, off-policy learning, and constrained MDPs. By leveraging contraction mappings, monotone operators, martingale techniques, mirror/proximal optimization, concentration inequalities, and mixing process theory, the study establishes finite-sample performance bounds and asymptotic convergence guarantees for diverse reinforcement learning algorithms, thereby forging a rigorous theoretical bridge between probability theory, optimization, and statistics.
This work addresses the uncertainty in payoff distributions arising from limited samples in data-driven games by proposing a distributionally robust game-theoretic framework grounded in coherent risk measures—such as Conditional Value-at-Risk and mean-semideviation—that internalize risk sensitivity as players’ preferences. It establishes, for the first time, a theoretical link between risk-awareness and distributional robustness. Leveraging multilinear complementarity programming and PPAD complexity analysis, the study proves the existence of equilibria under various ambiguity sets, characterizes their computational complexity, and quantifies the utility loss induced by risk aversion. Numerical experiments demonstrate that the proposed solutions exhibit superior out-of-sample performance and robustness.
This work addresses the degradation of conventional predictors trained on observational data when followers intervene on covariates to optimize their own objectives. Framing prediction and intervention as a Stackelberg game, the authors propose constructing robust predictors using invariant covariate sets—specifically, stable blankets. Theoretical analysis demonstrates that, under two common intervention objectives, stable blanket–based predictors match or outperform those based on causal parents and satisfy sufficient conditions for worst-case optimality. By integrating structural causal models, invariance-based learning, and graphical structure analysis, the method achieves a tighter theoretical risk bound and exhibits superior generalization performance, as validated on both synthetic and real-world datasets.