Score
Derives first-order and functional asymptotic approximations for transformations of estimators, producing analytic expressions for influence functions and asymptotic variances. Uses those approximations to propagate estimation uncertainty, compute variance estimates, and characterize efficiency bounds for estimators and statistical functionals.
This paper addresses the problem of estimating functionals of an unknown target function under a structure-agnostic setting—where no specific structural assumptions (e.g., Hölder smoothness) are imposed on the nuisance function, and only a generic convergence rate for nuisance estimation is assumed. Methodologically, it introduces the first formal framework for structure-agnostic estimation, operating under three simultaneous constraints: weak regularity conditions, compatibility with general-purpose nuisance estimators, and sample splitting. Theoretically, it establishes, for the first time, the essential optimality of first-order debiased estimators in this setting. Through minimax lower bound analysis, higher-order perturbation theory, and a unified debiasing framework, the paper precisely characterizes the optimal convergence rate and quantifies the fundamental trade-off between incorporating structural priors and improving estimation efficiency. These results provide foundational theoretical support for nonparametric and semiparametric inference.
This work addresses the lack of intuitive geometric interpretation in classical semiparametric efficiency theory, which has hindered the derivation and understanding of influence functions. The paper reformulates the theory within a differential geometric framework on the space of probability distributions, drawing an analogy to multivariate calculus: statistical paths, scores, and influence functions correspond respectively to curves, velocity vectors, and gradients. It demonstrates that the efficient influence function arises naturally as an orthogonal projection. By integrating functional analysis, differential geometry, and statistical inference, the study establishes a unified geometric interpretation of scores, tangent spaces, nuisance tangent spaces, and efficient influence functions. This synthesis not only clarifies several foundational theoretical issues but also substantially enhances the interpretability of methods in causal inference and missing data analysis.
This study addresses statistical inference for distributional models lacking classical density functions or finite moments by establishing a unified theoretical framework. The authors generalize Godambe’s inference functions to the space of distributions and introduce an observation operator to formally characterize diverse data-generating mechanisms, including point observations, interval censoring, and convolutional measurements. For the first time, inference functionals are defined in a distributional sense. Leveraging Schwartz’s theory of generalized functions, the Hájek–Le Cam convolution theorem, and Bhapkar–Godambe projections, the paper develops rigorous results on consistency, asymptotic normality, and optimality. A central contribution is the identification of a three-tier information hierarchy: Fisher information bounds the information captured by the observation operator, which in turn bounds the information attainable by any inference functional. The framework’s validity is demonstrated through applications to heavy-tailed distributions, interval-censored location models, and elliptical contour models.
Constructing efficient debiased estimators traditionally requires manual derivation of the efficient influence function (EIF), a labor-intensive process with high technical barriers and poor scalability. Method: This paper introduces Dimple, the first framework that models statistical functionals as compositions of differentiable primitives satisfying a novel differentiability condition; it leverages automatic differentiation to directly generate unbiased, efficient estimators while simultaneously identifying nuisance parameters. Dimple integrates probabilistic programming with functional decomposition, eliminating the need for explicit EIF derivation. Contribution/Results: We provide an open-source Python library enabling users to define parameters, generate estimators, and perform inference in just a few lines of code. Extensive experiments demonstrate Dimple’s effectiveness across diverse causal and semiparametric models—including AIPW, DR-Learner, and doubly robust IV—significantly lowering the barrier to constructing efficient estimators without sacrificing statistical efficiency.
This paper addresses complex causal problems—such as interference—that resist conventional experimental design, by proposing a unified functional-space framework for causal inference. Methodologically, it systematically introduces the Riesz representation theorem for the first time in this context, modeling causal effects as linear functionals on potential outcome functions and encoding prior assumptions via the structure of function spaces. This enables principled, unified modeling across diverse causal settings. Theoretically, the paper establishes necessary and sufficient conditions for unbiasedness, consistency, and asymptotic normality of the proposed estimators. Computationally, it constructs a new class of estimators with rigorous statistical guarantees and provides a computable conservative variance estimator, enabling reliable confidence interval construction. Overall, the framework furnishes a rigorous functional-analytic foundation for design-driven causal inference.
This study addresses the limitations of standard first-order semiparametric estimators in causal inference and missing data problems, which often fail to achieve asymptotic efficiency due to slow convergence of the nuisance functions and exhibit poor finite-sample performance. The authors systematically compare three classes of higher-order efficient estimators—Higher-Order Influence Functions (HOIF), kernel-based HOTMLE, and HAL-HOTMLE—evaluating, for the first time within a unified simulation framework, how their higher-order expansion constructions and regularization strategies affect estimation accuracy. Results demonstrate that higher-order debiasing substantially reduces bias, with HAL-HOTMLE showing robust performance, whereas HOIF proves sensitive to basis truncation and tuning parameters. The work clarifies the conditions under which higher-order corrections are effective in both theory and practice, while highlighting their limitations and key trade-offs for method selection.
This study addresses the estimation of parameters of the form θ₀ = E[F_Y⁻¹∘F_Z(X)] in the “changes-in-changes” model, for which existing methods lack theoretical guarantees when variables are unbounded. The authors construct a plug-in estimator based on empirical quantiles and establish its √n-consistency and asymptotic normality under assumptions weaker than those in the current literature. They further propose a novel consistent estimator for the asymptotic variance. The theoretical analysis leverages empirical process theory and plug-in methods for quantile functions. Monte Carlo simulations demonstrate that the proposed variance estimator substantially outperforms existing alternatives, leading to markedly improved inference accuracy.
This study addresses the challenge of characterizing the finite-sample distribution of ridge regression estimators, which hinders optimal regularization parameter selection and predictive performance. The authors propose a nonstandard asymptotic approach based on Gaussian approximation that accommodates heteroskedasticity and autocorrelation under a general data-generating mechanism. By introducing a local population parameter assumption and allowing the regularization parameter to vary with sample size, they establish the first effective finite-sample distributional approximation for low-dimensional ridge regression. Building on this approximation, they develop regularization parameter selection strategies that minimize either average or worst-case excess prediction risk, thereby substantially improving prediction accuracy.
This study investigates the asymptotic admissibility of Double Machine Learning (DML) estimators for quadratic functionals and integral functionals of densities under structural agnosticism. By integrating higher-order influence functions (HOIF), U-statistic theory, and a structure-free modeling framework, the authors establish—for the first time—that DML is asymptotically inadmissible for these two classes of functionals and construct a second-order influence function estimator that asymptotically dominates DML. For a third class of functionals, both DML and HOIF estimators achieve minimax optimality but neither dominates the other. These findings reveal fundamental limitations of DML under weak structural assumptions and provide superior alternatives grounded in higher-order influence functions.
This work addresses the computational and numerical challenges that commonly arise in practical implementations of higher-order influence function estimation, which often suffer from high-dimensional density estimation or inversion of large Gram matrices. The authors propose a stabilized estimation procedure that eliminates the need for sample splitting by incorporating Gram matrix regularization and a bilinear form structure. This approach avoids high-dimensional density estimation altogether while substantially improving numerical stability in finite samples. The method retains the theoretically optimal convergence rate and provides strong statistical guarantees alongside robust empirical performance, effectively overcoming key limitations of existing higher-order influence function estimators.