Adaptive inference for functionals of M-estimands

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the failure of classical inference under adaptive sampling and the reliance of existing methods on correct model specification. It proposes a unified framework featuring two novel approaches: a self-normalized statistic based on the quadratic variation of influence functions, and a plug-in estimator for reweighted incremental conditional variance. By leveraging Neyman orthogonality and nonparametric M-estimation theory, the framework constructs asymptotically valid confidence intervals for smooth functionals, accommodating flexible estimation while remaining robust to model misspecification. The effectiveness of these methods is demonstrated through dynamic pricing simulations, where they successfully produce asymptotically valid confidence intervals that standard approaches fail to generate.
πŸ“ Abstract
Reinforcement learning and contextual bandit algorithms have become increasingly common in sequential decision-making applications. When these methods are deployed in high-stakes domains, there is growing interest not only in learning effective policies, but also in conducting statistical inference for quantities learned under adaptive data collection. However, classical procedures applied naively in these settings can fail: even when estimators are unbiased, their variance becomes path-dependent and as a result may not be asymptotically normal. A growing literature has emerged to ameliorate this problem, but solutions tend to be problem specific and often rely on correct specification of a working model. In this work, we develop a unified framework for constructing asymptotically valid confidence intervals to cover smooth functionals of nonparametric M-estimands under adaptive sampling. Under Neyman orthogonality, we provide two novel methods for performing inference: (1) a self-normalized statistic based on the realized quadratic variation of the influence function and (2) a statistic using a plug-in estimate of the conditional variance based on reweighted influence function increments. Our results allow for flexible nonparametric estimation of nuisance parameters and remain valid under model misspecification. Our theory is supported by a simulation study for a dynamic pricing application which demonstrates that this method can produce asymptotically valid confidence intervals where standard methods fail.
Problem

Research questions and friction points this paper is trying to address.

adaptive inference
reinforcement learning
contextual bandits
M-estimands
statistical inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive inference
Neyman orthogonality
Self-normalized statistic
Nonparametric M-estimands
Contextual bandits