Predicting Long Term Sequential Policy Value Using Softer Surrogates

📅 2024-12-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional offline policy evaluation (OPE) struggles to rapidly and accurately estimate the long-term value of newly deployed policies—such as those in healthcare or education—when they introduce novel actions or operate in previously unseen environments. Method: We propose a transfer-based OPE framework leveraging short-horizon data. It integrates counterfactual reasoning, importance sampling, and doubly robust estimation to construct two novel estimators. For the first time, our approach enables testable, low-variance value transfer from short-horizon to full-horizon evaluation under policy distribution shift. Theoretical analysis establishes consistency and asymptotic normality. Results: Experiments on HIV and sepsis treatment simulators demonstrate that our method achieves statistically discriminative full-horizon value estimates using only 10% of the observational trajectory length—substantially outperforming existing OPE baselines.

Technology Category

Machine Learning: Online Learning & BanditsReasoning under Uncertainty: Sequential Decision MakingSearch and Optimization: Sampling/Simulation-based Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingResponsible Web: Human-perceived consequences of algorithmic deployment on the webUser Modeling, Personalization and Recommendation: Accountability, Transparency, and Ethics for personalization
📝 Abstract
Performing policy evaluation in education, healthcare and online commerce can be challenging, because it can require waiting substantial amounts of time to observe outcomes over the desired horizon of interest. While offline evaluation methods can be used to estimate the performance of a new decision policy from historical data in some cases, such methods struggle when the new policy involves novel actions or is being run in a new decision process with potentially different dynamics. Here we consider how to estimate the full-horizon value of a new decision policy using only short-horizon data from the new policy, and historical full-horizon data from a different behavior policy. We introduce two new estimators for this setting, including a doubly robust estimator, and provide formal analysis of their properties. Our empirical results on two realistic simulators, of HIV treatment and sepsis treatment, show that our methods can often provide informative estimates of a new decision policy ten times faster than waiting for the full horizon, highlighting that it may be possible to quickly identify if a new decision policy, involving new actions, is better or worse than existing past policies.
Problem

Research questions and friction points this paper is trying to address.

Policy Evaluation
Long-term Effects
Novel Policies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Policy Effect Prediction
Accelerated Forecasting
Data Integration Methodology
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Stanford University