Near-Equivalent Q-learning Policies for Dynamic Treatment Regimes

📅 2026-03-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes an ε-tolerance–based extension of Q-learning to address the limitation of conventional dynamic treatment regimes, which typically prescribe a single optimal policy while ignoring clinically relevant near-equivalent alternatives. By introducing a worst-value tolerance hyperparameter ε, the method generalizes the Q-function from a vector-valued to a matrix-valued representation, enabling the backward induction process to retain multiple acceptable value functions simultaneously. This formulation systematically captures the approximate equivalence inherent in treatment decisions. Integrating matrix-valued dynamic programming with a simulated tumor dynamics model, the approach successfully identifies regions of treatment indifference and generates sets of ε-optimal policies that offer meaningful clinical flexibility in both single- and multi-stage settings.

Technology Category

Planning, Routing, and Scheduling: Planning with Markov Models (MDPs, POMDPs)Reasoning under Uncertainty: Stochastic OptimizationSearch and Optimization: Sampling/Simulation-based Search

Application Category

Economics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystemsGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsUser Modeling, Personalization and Recommendation: User privacy protection in personalized systems
📝 Abstract
Precision medicine aims to tailor therapeutic decisions to individual patient characteristics. This objective is commonly formalized through dynamic treatment regimes, which use statistical and machine learning methods to derive sequential decision rules adapted to evolving clinical information. In most existing formulations, these approaches produce a single optimal treatment at each stage, leading to a unique decision sequence. However, in many clinical settings, several treatment options may yield similar expected outcomes, and focusing on a single optimal policy may conceal meaningful alternatives. We extend the Q-learning framework for retrospective data by introducing a worst-value tolerance criterion controlled by a hyperparameter $\varepsilon$, which specifies the maximum acceptable deviation from the optimal expected value. Rather than identifying a single optimal policy, the proposed approach constructs sets of $\varepsilon$-optimal policies whose performance remains within a controlled neighborhood of the optimum. This formulation shifts Q-learning from a vector-valued representation to a matrix-valued one, allowing multiple admissible value functions to coexist during backward recursion. The approach yields families of near-equivalent treatment strategies and explicitly identifies regions of treatment indifference where several decisions achieve comparable outcomes. We illustrate the framework in two settings: a single-stage problem highlighting indifference regions around the decision boundary, and a multi-stage decision process based on a simulated oncology model describing tumor size and treatment toxicity dynamics.
Problem

Research questions and friction points this paper is trying to address.

dynamic treatment regimes
near-equivalent policies
treatment indifference
Q-learning
precision medicine
Innovation

Methods, ideas, or system contributions that make the work stand out.

ε-optimal policies
Q-learning
dynamic treatment regimes
treatment indifference
matrix-valued Q-learning
🔎 Similar Papers
No similar papers found.
S
Sophia Yazzourh
Department of Epidemiology and Biostatistics, McGill University
E
Erica E. M. Moodie
Department of Epidemiology and Biostatistics, McGill University