🤖 AI Summary
This paper addresses the fundamental problem of quantifying, *a priori* and using only internal data, the potential efficiency gain conferred by external information on statistical estimation—without access to or reliance on that external information itself. We propose a model- and method-agnostic utility metric—the *estimation efficiency ratio*—grounded in the semiparametric efficiency bound, which characterizes the maximal theoretical improvement attainable from external information. Leveraging the efficient influence function, we construct a computationally feasible estimator solely from internal data and establish its asymptotic normality, enabling both point estimation and valid confidence interval inference. Simulation studies and empirical applications demonstrate robust finite-sample performance. The framework provides a theoretically optimal, broadly applicable pre-evaluation tool for cost-sensitive decisions such as data acquisition and collaborative modeling.
📝 Abstract
We consider a general statistical estimation problem involving a finite-dimensional target parameter vector. Beyond an internal data set drawn from the population distribution, external information, such as additional individual data or summary statistics, can potentially improve the estimation when incorporated via appropriate data fusion techniques. However, since acquiring external information often incurs costs, it is desirable to assess its utility beforehand using only the internal data. To address this need, we introduce a utility measure based on estimation efficiency, defined as the ratio of semiparametric efficiency bounds for estimating the target parameters with versus without incorporating the external information. It quantifies the maximum potential efficiency improvement offered by the external information, independent of specific estimation methods. To enable inference on this measure before acquiring the external information, we propose a general approach for constructing its estimators using only the internal data, adopting the efficient influence function methodology. Several concrete examples, where the target parameters and external information take various forms, are explored, demonstrating the versatility of our general framework. For each example, we construct point and interval estimators for the proposed measure and establish their asymptotic properties. Simulation studies confirm the finite-sample performance of our approach, while a real data application highlights its practical value. In scientific research and business applications, our framework significantly empowers cost-effective decision making regarding acquisition of external information.