stochastic policy design

Design, parameterize, and evaluate stochastic policies that map states or covariates to probability distributions — including function-valued or distributional interventions — by building implementable parametric families or basis representations for policy shifts and constructing estimands and estimators that may avoid reliance on strict positivity. In multi‑agent or adaptive settings, formulate and analyze game‑theoretic reinforcement‑learning and control models to study incentive alignment, strategic equilibria, and the stability of interacting stochastic policies.

stochasticpolicydesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Policy Learning with Distributional Welfare

Nov 27, 2023
YC
Yifan Cui
🏛️ Zhejiang University | University of Bristol

This paper addresses the vulnerability of conventional utilitarian policy learning—based on the conditional average treatment effect (CATE)—to outliers and its inability to flexibly accommodate policy caution or leniency under heterogeneous individual treatment effects. We propose an optimal intervention allocation framework grounded in the conditional quantile treatment effect (QoTE). To our knowledge, this is the first work to incorporate distributionally robust welfare into policy learning, formulating a minimax strategy based on QoTE that accommodates the fundamental challenge of non-point-identification of the counterfactual joint distribution. Our method integrates causal inference, distributionally robust optimization, and decision theory, supporting both stochastic and deterministic policies under various identification assumptions. We establish an asymptotically tight upper bound on the regret and demonstrate robustness to model misspecification. The framework generalizes to any welfare objective defined as a functional of the potential outcomes’ joint distribution.

Generalizing welfare frameworks beyond utilitarian criteria for heterogeneous populationsIdentifying policies robust to uncertainty in counterfactual outcome distributionsOptimal treatment allocation targeting distributional welfare, not just average effects

Statistical Learning of Distributionally Robust Stochastic Control in Continuous State Spaces

Jun 17, 2024
SW
Shengbo Wang
🏛️ USC | HKUST | Stanford University | New York University

This paper addresses distributionally robust stochastic control in continuous state spaces to mitigate the policy fragility of conventional i.i.d. Markov models—arising from neglecting environmental input distribution shifts and endogenous dependencies. We propose a novel paradigm that balances modeling simplicity and robustness: adaptive adversarial perturbations are embedded within dynamic programming to unify *f*-divergence and Wasserstein-type ambiguity sets. We establish, for the first time in continuous spaces, a unified learning theory for robust value functions under both ambiguity-set classes, providing finite-sample minimax convergence rate bounds. Integrating distributionally robust optimization, stochastic control, and nonparametric statistics, we design a computationally tractable minimax policy learning algorithm. Experiments demonstrate that the framework achieves both statistical efficiency and strong robustness across real-world applications—including supply chain management and finance.

Addressing policy fragility through distributionally robust adversarial perturbation modelsDeveloping minimax statistical rates and deep RL algorithms for robust policy optimizationLearning robust stochastic control for continuous state-action spaces with data-driven approaches

Asymptotic Theory for IV-Based Reinforcement Learning with Potential Endogeneity

Mar 06, 2021
JL
Jin Li
🏛️ The University of Hong Kong | The Hong Kong University of Science and Technology

This paper addresses the reinforcement bias and endogeneity arising from the dynamic coupling between data generation and policy evaluation in reinforcement learning (RL). We propose Instrumental Variable RL (IV-RL), the first RL paradigm explicitly designed to handle endogeneity via instrumental variables. By modeling policy iteration as a Markov process with instrument-dependent transitions, we develop an asymptotic statistical theory for IV-RL, derive an inferentially valid estimator of the optimal policy, and quantify how temporal dependence degrades inference accuracy. Our method integrates instrumental variable estimation, stochastic approximation theory, and Markov decision process (MDP) modeling. We rigorously establish strong consistency and asymptotic normality of the IV-RL estimator, enabling unbiased policy evaluation and principled confidence interval construction. The key contribution is breaking the conventional offline RL assumption of exogenous data: IV-RL provides the first statistically grounded, endogeneity-robust correction framework for dynamic decision-making under endogenous feedback.

Data AnalysisMarkov Decision ProcessReinforcement Bias

Logit-Q Dynamics for Efficient Learning in Stochastic Teams

Feb 20, 2023
AS
Ahmed Said Donmez
🏛️ Cornell University | Bilkent University

This work addresses stochastic team games under unknown dynamics. Methodologically, it proposes the Logit-Q dynamics framework—the first to couple Logit-response dynamics with Q-learning within an auxiliary stage game—where Q-functions drive state-dependent payoffs to enable efficient equilibrium learning. Technically, it introduces a novel analytical approach combining fictitious static Q-estimation scenarios with asymptotic coupling to the true dynamic environment, integrated with slowly varying epoch scheduling and coupling-based convergence analysis. This yields the first convergence and rationality guarantees for non-fully controllable stochastic games. Theoretically, the algorithm converges to an approximately optimal team equilibrium, with quantifiable approximation error; exhibits rationality against pure stationary-strategy opponents; and retains convergence when stage payoffs form a potential game and state transitions are controlled by a single agent.

Achieves near-efficient equilibrium in teams with unknown dynamics.Develops logit-Q dynamics for efficient learning in stochastic games.Ensures convergence in games with potential stage-payoffs.

Post Reinforcement Learning Inference

Feb 17, 2023
VS
Vasilis Syrgkanis
🏛️ Stanford University | Hong Kong University of Science and Technology

In reinforcement learning, adaptive interaction data—where the behavior policy is nonstationary—invalidates standard estimators, undermining asymptotic normality for off-policy counterfactual policy evaluation and dynamic treatment effect (DTE) inference. To address this, we propose a weighted Z-estimation framework that constructs time-varying adaptive weights to stabilize heteroskedasticity, achieving, for the first time in the RL off-policy setting, both consistent and asymptotically normal DTE estimation. Our approach integrates dynamic causal inference with asymptotic statistical theory, enabling rigorous hypothesis testing and construction of uniformly valid confidence regions. Simulation studies and real-world RL experiments demonstrate substantial improvements in confidence interval coverage and statistical power. The method provides the first solution for structural parameter inference under adaptive experimentation that simultaneously offers theoretical guarantees—namely consistency, asymptotic normality, and uniform validity—and empirical robustness.

Address nonstationary variance in adaptive reinforcement learning environmentsDevelop weighted Z-estimation for dynamic treatment effect analysisEstimate counterfactual policies post reinforcement learning data collection

Latest Papers

What's happening recently
View more

This work addresses the lack of a unified theoretical framework for reinforcement learning, which has hindered systematic analysis of its convergence, sample complexity, and generalization. Building upon Markov decision processes and Bellman operators, the paper introduces a cohesive analytical framework that integrates tools from operator theory, stochastic approximation, convex duality, and function approximation. This framework encompasses a broad range of algorithms, including value iteration, policy iteration, temporal difference methods, off-policy learning, and constrained MDPs. By leveraging contraction mappings, monotone operators, martingale techniques, mirror/proximal optimization, concentration inequalities, and mixing process theory, the study establishes finite-sample performance bounds and asymptotic convergence guarantees for diverse reinforcement learning algorithms, thereby forging a rigorous theoretical bridge between probability theory, optimization, and statistics.

function approximationMarkov decision processesmathematical foundations

This work addresses the uncertainty in payoff distributions arising from limited samples in data-driven games by proposing a distributionally robust game-theoretic framework grounded in coherent risk measures—such as Conditional Value-at-Risk and mean-semideviation—that internalize risk sensitivity as players’ preferences. It establishes, for the first time, a theoretical link between risk-awareness and distributional robustness. Leveraging multilinear complementarity programming and PPAD complexity analysis, the study proves the existence of equilibria under various ambiguity sets, characterizes their computational complexity, and quantifies the utility loss induced by risk aversion. Numerical experiments demonstrate that the proposed solutions exhibit superior out-of-sample performance and robustness.

coherent risk measuresdistributional uncertaintydistributionally robust games

This work addresses the degradation of conventional predictors trained on observational data when followers intervene on covariates to optimize their own objectives. Framing prediction and intervention as a Stackelberg game, the authors propose constructing robust predictors using invariant covariate sets—specifically, stable blankets. Theoretical analysis demonstrates that, under two common intervention objectives, stable blanket–based predictors match or outperform those based on causal parents and satisfy sufficient conditions for worst-case optimality. By integrating structural causal models, invariance-based learning, and graphical structure analysis, the method achieves a tighter theoretical risk bound and exhibits superior generalization performance, as validated on both synthetic and real-world datasets.

causal inferencedistributional robustnessinvariant sets

Hot Scholars

NP

Nikolaos Pappas

Associate Professor, Linköping University (LiU)
Age of InformationSemantic CommunicationsCommunication Networks
XE

Xin Eric Wang

Assistant Professor, University of California, Santa Barbara, Simular
NLPCVMLLanguage and Vision
JP

Jan Peters

Professor for Intelligent Autonomous Systems/TU Darmstadt, Dept. Head/German AI Research Center DFKI
Robot LearningReinforcement LearningMachine LearningRobotics
GF

Gabriele Farina

Assistant Professor of Computer Science, MIT
Computational Game TheoryOptimizationEconomics and Computation
UT

Ufuk Topcu

The University of Texas at Austin
autonomycontrolsformal methodslearning