Score
Design, build, or analyze algorithms that learn sequential, personalized treatment policies online using reinforcement-learning and online policy-learning methods. Work includes continual updating from incoming cases, safe evaluation and policy improvement via simulated rollouts, safety-constrained or rule-aware policy optimization, and mechanisms to allow clinician overrides.
Online platforms often discretize continuous incentives for A/B testing, which hinders extrapolation to untested intervention levels and overlooks user heterogeneity, leading to suboptimal decisions. This work proposes a Deep Learning for Policy Targeting (DLPT) framework that establishes, for the first time, a theoretical foundation for learning personalized continuous intervention policies from discrete randomized controlled trials. We prove the asymptotic unbiasedness and consistency of the policy value estimator and derive a root-n regret bound. By integrating high-dimensional user features, DLPT enables end-to-end optimization of personalized continuous policies. In a real-world incentive experiment conducted in collaboration with a leading social media platform, DLPT substantially outperforms existing benchmarks, achieving significant improvements in both policy value estimation and identification of personalized optimal interventions.
Learning continuous treatment policies from observational data faces three key challenges: nonparametric welfare estimation, infinite-dimensional policy spaces, and shape constraints (e.g., monotonicity or convexity). Method: We propose a novel paradigm that approximates the shape-constrained policy space via a sequence of finite-dimensional subspaces; we develop a data-adaptive tuned penalization algorithm, integrating kernel-based welfare estimation, machine learning–based propensity score modeling, regularized optimization, and constrained function approximation. Contribution/Results: We establish, for the first time, an oracle inequality for welfare regret under continuous treatments. Theoretically, our estimator achieves statistically optimal convergence rates under both known and unknown propensity scores. Empirically, it significantly enhances out-of-sample policy extrapolation robustness and real-world effectiveness.
Offline pretraining often suffers from rapid degradation and poor exploration during early online reinforcement learning. To address this, we propose a policy expansion mechanism that treats the frozen offline policy as a fixed behavioral prior, dynamically coordinating it with a learnable online policy. Our key contribution is the first adaptive dual-policy architecture, integrating a policy-ensemble-based gating mechanism with behavioral distribution matching constraints. This ensures the offline policy remains unupdated while continuously guiding exploration, while enabling the online policy to incrementally acquire novel behaviors. Evaluated on multiple continuous-control benchmarks, our method significantly improves sample efficiency and final performance, avoids initial performance collapse, and achieves more stable convergence—outperforming standard fine-tuning and policy distillation baselines across all metrics.
Deterministic algorithms used in U.S. pretrial risk assessment lack strategic optimizability and interpretability, limiting their ability to improve upon existing policies safely and transparently. Method: We propose an extrapolation-based safe policy learning framework that integrates robust optimization (maximizing minimum expected utility), partial identification theory, and causal inference—enabling safe policy improvement without assuming randomized interventions. Contribution/Results: To our knowledge, this is the first method to guarantee, with high probability, that a new deterministic policy dominates the current one under observational data, while ensuring statistical safety and interpretability. Evaluated on real-world field experiment data, our approach significantly increases the proportion of defendants in specific subgroups who are safely downgraded to “low risk,” without compromising overall decision safety. This advances responsible deployment of judicial algorithms by providing a principled, auditable, and policy-improving paradigm.
This paper addresses optimal policy assignment under partial treatment coverage in heterogeneous populations, where experiments only implement a subset of possible treatment values, limiting generalizability to untested interventions. Method: We propose the first framework integrating shape-constrained partial identification of treatment effects—incorporating monotonicity or convexity constraints—with a minimax regret criterion, formulated as a tractable mixed-integer linear program (MILP). The method combines nonparametric conditional average treatment effect (CATE) estimation, shape restrictions, minimax optimization, and efficient linear/integer programming solvers. Contribution/Results: Applied to a Kenyan rural electricity subsidy experiment, our framework recommends novel, experimentally untested treatment levels—covering nearly the entire population—while reducing maximum regret by over 60%. It substantially enhances policy extrapolation capability and robustness beyond conventional methods constrained to observed treatment supports.
This work addresses the challenge of dynamic medical treatment, which requires joint optimization of treatment intensity and interaction timing. Existing approaches often rely on fixed interaction intervals or enforce safety only at discrete time points, failing to account for continuous state evolution and intermediate risks. The authors formulate the problem as an options-based semi-Markov decision process with trajectory-level safety constraints, where each option comprises a continuous-time treatment policy and its duration. Key contributions include a safety tightening mechanism that provably ensures trajectory-wide safety with high probability by imposing appropriate constraints at interaction times, a finite-sample policy learning theory grounded in logged data, and a data-driven conservative surrogate method. Experiments demonstrate that the proposed adaptive interaction mechanism significantly outperforms fixed-interval strategies across multiple safety policies, enhancing both treatment safety and efficacy.
This work proposes an ε-tolerance–based extension of Q-learning to address the limitation of conventional dynamic treatment regimes, which typically prescribe a single optimal policy while ignoring clinically relevant near-equivalent alternatives. By introducing a worst-value tolerance hyperparameter ε, the method generalizes the Q-function from a vector-valued to a matrix-valued representation, enabling the backward induction process to retain multiple acceptable value functions simultaneously. This formulation systematically captures the approximate equivalence inherent in treatment decisions. Integrating matrix-valued dynamic programming with a simulated tumor dynamics model, the approach successfully identifies regions of treatment indifference and generates sets of ε-optimal policies that offer meaningful clinical flexibility in both single- and multi-stage settings.
This study addresses the challenge of dynamic, personalized decision-making in recurrent bladder cancer treatment, a task inadequately handled by existing systems that rely on static guidelines or single-step predictions and fail to model the temporal evolution of disease. To overcome this limitation, the authors propose the first decision support framework integrating recurrent patient state transition simulation with deep reinforcement learning. Built upon a Markov decision process and a deep Q-network, the framework enables end-to-end training to generate interpretable, temporally coherent treatment plans alongside transparent decision logs. Experimental results in a simulated clinical environment demonstrate the approach’s efficacy and robustness, achieving a cumulative reward of 63,918.87, an average training loss of 0.0056, and a policy improvement rate of 6.62%.
This paper addresses safety risks in high-stakes personalized decision-making arising from policy learning models that are forced to produce outputs regardless of prediction uncertainty. To mitigate this, we propose a *policy learning framework with abstention*, enabling the model to deliberately abstain—deferring to a safe default policy or expert intervention—when predictive uncertainty is high. Methodologically, we design a two-stage learning framework: first, we construct an abstention rule based on disagreement among approximately optimal policies; second, we extend it to marginal conditional modeling, distributionally robust optimization, and safe policy improvement. We employ a doubly robust objective to handle unknown propensity scores and incorporate an O(1/n) regret bound analysis alongside a stochastic reward compensation mechanism. Theoretically, we establish fast-converging regret bounds under both known and unknown propensity scores. Our approach significantly enhances the safety, robustness, and practicality of policy learning in critical applications.
This work addresses the challenge of efficiently evaluating personalized sequential therapies for Alzheimer’s disease (AD), which is hindered by the disease’s prolonged progression and high patient heterogeneity. To this end, the authors propose ALPACA—an open-source, Gym-compatible reinforcement learning environment that integrates a Conditional Autoregressive State Transition (CAST) model with reinforcement learning, leveraging longitudinal data from the Alzheimer’s Disease Neuroimaging Initiative (ADNI). ALPACA constitutes the first interpretable and reusable digital trial platform for AD, enabling drug repurposing exploration and individualized treatment optimization while generating clinically plausible conditional treatment trajectories. Reinforcement learning policies trained within ALPACA significantly outperform both no-treatment and behavior cloning baselines on memory-related outcomes, with their decision-making grounded in clinically meaningful patient features.