🤖 AI Summary
This study addresses the high-variance problem in off-policy evaluation caused by insufficient behavior policy coverage by investigating multi-source counterfactual annotation allocation strategies for contextual bandits under budget constraints. We propose an integer programming allocation model coupled with a majorization-minimization (MM) optimization algorithm and dynamic programming subroutines. This approach precisely characterizes annotation value thresholds and local mechanisms, thereby achieving optimal estimator configuration. Experiments on synthetic clinical and educational datasets demonstrate that the proposed method reduces mean squared error by 20.58% and 10.77%, respectively, effectively enhancing evaluation accuracy under limited budgets.
📝 Abstract
Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation. Counterfactual annotations can add evidence about unobserved actions, yet practical sources, including domain experts and large language models (LLMs), may be costly, biased, or noisy. We study budgeted acquisition of such annotations for contextual-bandit OPE. Given source-specific costs and error profiles, we formulate an integer allocation problem over context-action pairs and annotation sources to minimize the component of estimator variance that depends on the annotation plan. We characterize when annotations are valuable through a first-annotation threshold and local annotation-value regimes. For the coupled multi-source problem, we develop a majorization-minimization algorithm with dynamic-programming subroutines that monotonically improves the objective. Experiments in synthetic clinical and LLM-annotated education bandits show that our allocation method reduces fixed-profile mean squared error (MSE) by 20.58% and 10.77%, respectively, relative to no annotation.