Score
Designs and analyzes estimators and algorithms that quantify how removing or perturbing a single observation, resource, or unit affects future predictions or downstream metrics (forward or future influence). This work builds and evaluates influence-function linearizations, higher-order influence (HOIF) and second-order U-statistic corrections, leave-one-out/leave-one-resource and ablation procedures, empirical HOIF estimators, per-instance or per-document influence measures, and the semiparametric/efficient-inference tools needed for uncertainty quantification and hypothesis testing.
This study addresses the limitations of standard first-order semiparametric estimators in causal inference and missing data problems, which often fail to achieve asymptotic efficiency due to slow convergence of the nuisance functions and exhibit poor finite-sample performance. The authors systematically compare three classes of higher-order efficient estimators—Higher-Order Influence Functions (HOIF), kernel-based HOTMLE, and HAL-HOTMLE—evaluating, for the first time within a unified simulation framework, how their higher-order expansion constructions and regularization strategies affect estimation accuracy. Results demonstrate that higher-order debiasing substantially reduces bias, with HAL-HOTMLE showing robust performance, whereas HOIF proves sensitive to basis truncation and tuning parameters. The work clarifies the conditions under which higher-order corrections are effective in both theory and practice, while highlighting their limitations and key trade-offs for method selection.
This work addresses the suboptimal convergence rates of classical influence functions when estimating complex, implicitly defined causal parameters such as quantile treatment effects. While existing higher-order methods are limited to explicitly defined parameters, this paper extends the higher-order influence function framework to implicit M- and Z-estimation problems for the first time. By integrating U-process theory with nonparametric estimation, the authors construct a debiased estimator that substantially relaxes the stringent Hölder smoothness assumptions typically imposed on nuisance parameters. The proposed approach achieves improved convergence rates in settings like quantile treatment effect estimation and reduces requirements on model complexity, thereby broadening the applicability of higher-order influence function methodology to a wider class of semiparametric problems.
To address the low precision of average treatment effect (ATE) estimation in randomized controlled trials (RCTs) with high-dimensional covariates (p ≫ n), this paper proposes a novel covariate adjustment method based on higher-order influence functions (HOIFs). The method systematically establishes the theoretical advantages of HOIFs in RCTs for the first time, unifies a broad class of state-of-the-art adjusted estimators, and rigorously characterizes the conditions under which HOIF-based estimation strictly dominates both unadjusted and linear-model-adjusted estimators. We prove that the proposed estimator achieves semiparametric asymptotic efficiency—i.e., it attains the semiparametric efficiency bound under mild regularity conditions. Numerical simulations and empirical analyses demonstrate substantial gains in estimation accuracy and robustness when p is large relative to n. An accompanying R package, implementing the method, has been publicly released on CRAN.
This work investigates the impact of infinitesimal perturbations to training data on model performance, aiming to enhance model interpretability and robustness. Addressing the high computational cost and limited interpretability of conventional influence functions—particularly in non-convex models—the paper pioneers the integration of Fisher information geometry into influence estimation. It introduces the Approximate Fisher Influence Function (AFIF), reformulating influence estimation as a weighted empirical risk minimization problem. Leveraging information-geometric principles, the authors derive an efficient approximation algorithm that avoids explicit Hessian inversion. The method achieves several-fold speedup over Newton-type approaches while maintaining high accuracy and strong robustness across both generalized linear models and non-convex neural networks. By unifying geometric insight with practical scalability, AFIF bridges interpretability and usability in influence analysis.
This paper studies networked innovation processes, where each process is modeled as an infinite-color Pólya urn to capture novelty emergence. Addressing the limitation of existing models—which ignore historical interdependencies among processes—we develop, for the first time, a second-order asymptotic theory for interactive innovation processes, characterizing their joint growth rates and covariance structure. We propose a general statistical framework based on intensity function estimation and point-process inference to quantify the direction and magnitude of cross-process influence. The methodology is empirically validated on Reddit community evolution and Gutenberg textual innovation data, demonstrating both theoretical consistency and statistical robustness. Our approach provides a scalable theoretical toolkit and practical methodology for modeling innovation diffusion and conducting causal inference across diverse domains.
In causal effect estimation, the absence of standardized hyperparameter tuning evaluation criteria impedes reliable model selection and creates a substantial gap between commonly used metrics and true performance. This paper systematically investigates the interplay between hyperparameter tuning and evaluation, jointly analyzing estimators (T-/X-/R-Learner), base learners (random forests, gradient boosting, neural networks), and evaluation metrics (IPW, DR, PEHE) across four benchmark datasets. Key findings are: (1) thorough hyperparameter tuning eliminates performance differences among mainstream causal estimators; (2) the choice of evaluation strategy exerts greater influence on final performance than either the estimator type or base learner architecture; and (3) existing evaluation metrics underestimate the performance gain from optimal model selection by over 35% on average. These results demonstrate that hyperparameter tuning is the primary determinant of causal estimation accuracy, underscoring an urgent need for more robust, theoretically grounded evaluation paradigms in causal machine learning.
This study addresses the long-standing misconception that estimator ranking inconsistencies in data attribution stem from approximation errors, revealing instead that they originate from counterfactual norm mismatches. By formalizing influence as a counterfactual estimator, this work establishes norm analysis as a necessary prerequisite for comparing estimators. It derives local decompositions to analytically characterize signal interaction mechanisms, validated through linearized approximations and controlled experiments. The primary contribution is the first demonstration that behavioral proxy selection critically impacts attribution quality, proving that differing norms directly induce ranking discrepancies. Furthermore, the proposed behavior-aligned norm successfully identifies target samples overlooked by default methods, substantially improving attribution accuracy.
This work addresses the computational and numerical challenges that commonly arise in practical implementations of higher-order influence function estimation, which often suffer from high-dimensional density estimation or inversion of large Gram matrices. The authors propose a stabilized estimation procedure that eliminates the need for sample splitting by incorporating Gram matrix regularization and a bilinear form structure. This approach avoids high-dimensional density estimation altogether while substantially improving numerical stability in finite samples. The method retains the theoretically optimal convergence rate and provides strong statistical guarantees alongside robust empirical performance, effectively overcoming key limitations of existing higher-order influence function estimators.
This work addresses the lack of intuitive geometric interpretation in classical semiparametric efficiency theory, which has hindered the derivation and understanding of influence functions. The paper reformulates the theory within a differential geometric framework on the space of probability distributions, drawing an analogy to multivariate calculus: statistical paths, scores, and influence functions correspond respectively to curves, velocity vectors, and gradients. It demonstrates that the efficient influence function arises naturally as an orthogonal projection. By integrating functional analysis, differential geometry, and statistical inference, the study establishes a unified geometric interpretation of scores, tangent spaces, nuisance tangent spaces, and efficient influence functions. This synthesis not only clarifies several foundational theoretical issues but also substantially enhances the interpretability of methods in causal inference and missing data analysis.
This work addresses the lack of precise characterization of individual sample influence in high-dimensional convex M-estimation when the sample size and dimensionality are of the same order (\(n \sim d\)). Under Gaussian design assumptions, the authors rigorously derive the asymptotic distribution of influence measures by integrating high-dimensional asymptotic analysis, random matrix theory, and leave-one-out influence diagnostics. They establish that the influence measure converges to an explicitly expressible limiting distribution. The analysis further reveals that highly influential samples concentrate near the decision boundary, thereby providing theoretical justification for boundary-based heuristic strategies in active learning and uncovering an intrinsic connection between influence analysis and classification boundaries.
This study addresses the challenge of detecting strongly influential outliers in clustered data within mixed-effects models, where existing methods are often constrained by model assumptions and lack generalizable diagnostic tools. To overcome these limitations, this work proposes a model-agnostic, point-level anomaly detection framework that integrates SHAP values with residuals to construct comprehensive features. It leverages Normalizing Flows to map complex distributions into a standard space, enabling precise outlier identification across diverse base learners, including linear models, random forests, and gradient boosting trees. The primary contributions include providing goodness-of-fit diagnostics through flow-based modeling, empirically validating the method’s effectiveness and generalizability across multiple architectures, and systematically analyzing its advantages and limitations.