🤖 AI Summary
This study addresses the challenge that existing expected value of sample information (EVSI) computation methods struggle to accommodate missing data, a common issue in real-world research. The authors propose a novel EVSI estimation approach tailored for complex missingness mechanisms—including missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR)—by integrating individual-level data simulation, multiple imputation, and nonparametric regression. Notably, this work extends the EVSI framework to settings with non-ignorable missingness for the first time. Building on the EVSI loss, the study further introduces a new sample size determination strategy. Findings demonstrate that missing data substantially reduce EVSI, and achieving the EVSI attainable under complete data requires markedly larger sample sizes than conventional calculations suggest, thereby enhancing the practical applicability of health economic decision models.
📝 Abstract
The Expected Value of Sample Information (EVSI) is a powerful instrument to determine the value of additional evidence to inform an economic model. However, EVSI has been applied only to idealized data collection mechanisms, thereby reducing its potential applications in realistic studies. In this paper, we define a methodology to calculate EVSI when the additional evidence we aim to collect exhibits missing data; a very common challenge in real-world studies. First, we define how to simulate individual-level data and how to induce missingness inside the simulated data. We will reproduce Missing Completely At Random (MCAR), Missing At Random (MAR), and Missing Not At Random (MNAR) missing data mechanisms. Then, we will apply the multiple imputation method to adjust for the bias related to the missing data. Finally, we use the imputed data to compute the EVSI using nonparametric regression methods. We apply the novel methodology to two different health economic models and compare the EVSI computed on data without missingness with the EVSI with different types of missing data. We show that the EVSI decreases when the additional evidence suffers from missingness, and therefore, we define a method to efficiently compute the sample size we need to collect to recover the idealized EVSI (without missingness). We find out that the number of additional samples needed to correct for missing data exceeds that coming from standard approaches. With this methodology, we compute EVSI when the additional data are affected by non-trivial forms of missingness, modeling both the MAR and MNAR mechanisms, and extending EVSI calculation to more realistic scenarios. The fact that the necessary sample size to recover the idealized EVSI exceeds that of standard methods suggests that this methodology potentially represents a novel technique for sample size calculation in realistic studies.