Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling

πŸ“… 2026-10-06
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the problem of excessive action-level weight variance in off-policy evaluation for contextual multi-armed bandits by proposing the VOCEM estimator. This method constructs an unbiased interpolation family between the OffCEM and doubly robust estimators, deriving a closed-form optimal interpolation coefficient to minimize estimation variance while theoretically guaranteeing that the resulting variance is strictly lower than that of both endpoint estimators. By integrating joint effect modeling with variance optimization theory, VOCEM achieves significant improvements over existing baselines across 23 experimental configurations on both synthetic and real-world benchmarks. These results demonstrate that the proposed approach effectively enhances the stability and robustness of off-policy evaluation.
πŸ“ Abstract
Off-policy evaluation (OPE) for contextual bandit policies becomes challenging when action-level importance weighting incurs excessive variance. Doubly robust (DR) estimation remains unbiased under common support but retains these high-variance action-level weights. A prior estimator, Off-policy evaluation with Conjunct Effect Model (OffCEM), replaces them with more stable cluster-level weights, at the cost of relying on local correctness of the reward model. In this paper, we show that, under the assumptions required by DR and OffCEM, there exists an unbiased family of estimators that interpolates between OffCEM and DR. Building on this result, we propose the Variance Optimal-CEM (VOCEM) estimator, which selects the interpolation coefficient to minimize variance. We derive the population-optimal coefficient in closed form and show that the resulting estimator has variance no larger than either endpoint, OffCEM or DR. Experiments in controlled synthetic settings and on two large-action benchmarks show that VOCEM improves upon both endpoints in all 23 evaluated conditions, exhibiting greater stability and empirical robustness.
Problem

Research questions and friction points this paper is trying to address.

off-policy evaluation
contextual bandits
variance reduction
importance weighting
doubly robust estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Off-Policy Evaluation
Doubly Robust Estimation
Variance Optimization
Conjunct Effect Model
Contextual Bandits
πŸ”Ž Similar Papers
No similar papers found.