Score
Designs and computes influence functions and efficient influence functions (EIFs) for estimators by representing perturbations as distributional paths and deriving pathwise derivatives and orthogonal projections, implementing algorithms to compute influence-function scores and sensitivity matrices. Builds local influence diagnostics for case-weight and response perturbation analysis to quantify per-sample training effects, detect influential or poisonous observations, and evaluate estimator robustness via curvature and sensitivity measures.
This work investigates the impact of infinitesimal perturbations to training data on model performance, aiming to enhance model interpretability and robustness. Addressing the high computational cost and limited interpretability of conventional influence functions—particularly in non-convex models—the paper pioneers the integration of Fisher information geometry into influence estimation. It introduces the Approximate Fisher Influence Function (AFIF), reformulating influence estimation as a weighted empirical risk minimization problem. Leveraging information-geometric principles, the authors derive an efficient approximation algorithm that avoids explicit Hessian inversion. The method achieves several-fold speedup over Newton-type approaches while maintaining high accuracy and strong robustness across both generalized linear models and non-convex neural networks. By unifying geometric insight with practical scalability, AFIF bridges interpretability and usability in influence analysis.
In causal inference, machine learning is widely used to estimate propensity scores and outcome models, yet practical guidance for constructing target minimum loss estimators (TMLEs) that balance theoretical rigor with implementability remains scarce for applied researchers. Method: Building on the efficient influence function framework, we propose a modular, reproducible TMLE construction procedure: first fit auxiliary models using arbitrary machine learning methods; then perform a one-dimensional targeted update to correct bias. Contribution/Results: Our key innovation lies in translating abstract efficiency theory into three intuitive steps—initial estimation, efficient influence function computation, and parameter update—substantially lowering the conceptual and implementation barriers to TMLE. The estimator retains double robustness and asymptotic efficiency under model misspecification. By bridging the critical gap between statistical theory and empirical practice, our approach enables non-statisticians to conduct reliable, machine learning–enhanced causal analyses.
In causal mediation analysis, conventional estimators of the mediated effect functional suffer from low accuracy and high sensitivity to misspecification of nuisance functions. To address this, we propose a bias-structure-guided two-stage framework that decouples nuisance function estimation. In Stage I, we estimate only the bias-relevant component of the mediation mechanism—rather than the full mechanism—thereby reducing model dependence. In Stage II, we introduce a nonparametric weighted balancing estimator, where weights are constructed by directly optimizing the asymptotic bias of the mediated effect estimator. We establish theoretical guarantees: the resulting estimator is consistent and asymptotically normal, and remains robust under partial misspecification of nuisance functions. Compared with standard approaches, our method substantially improves estimation accuracy and reliability. It provides a principled tool for mediation inference in high-dimensional settings or under model uncertainty.
Constructing efficient debiased estimators traditionally requires manual derivation of the efficient influence function (EIF), a labor-intensive process with high technical barriers and poor scalability. Method: This paper introduces Dimple, the first framework that models statistical functionals as compositions of differentiable primitives satisfying a novel differentiability condition; it leverages automatic differentiation to directly generate unbiased, efficient estimators while simultaneously identifying nuisance parameters. Dimple integrates probabilistic programming with functional decomposition, eliminating the need for explicit EIF derivation. Contribution/Results: We provide an open-source Python library enabling users to define parameters, generate estimators, and perform inference in just a few lines of code. Extensive experiments demonstrate Dimple’s effectiveness across diverse causal and semiparametric models—including AIPW, DR-Learner, and doubly robust IV—significantly lowering the barrier to constructing efficient estimators without sacrificing statistical efficiency.
Quantifying channel importance and performing channel-centric analysis remain challenging in multivariate time series (MTS) modeling. Method: This paper introduces the Channel-level Influence Function (CIF), the first influence-function-based approach tailored for MTS, grounded in robust statistics theory. CIF leverages first-order gradient approximations and dataset-averaged gradients to estimate the causal contribution of each input channel to model outputs. Unlike conventional global influence functions, CIF enables fine-grained, interpretable channel importance assessment and supports principled channel pruning. Results: Extensive evaluation across multiple real-world MTS datasets demonstrates that CIF significantly improves anomaly detection and forecasting accuracy over baseline methods. Moreover, CIF is the only existing framework capable of enabling stable, performance-preserving channel pruning while maintaining model robustness—establishing it as a novel, theoretically grounded quantification paradigm for MTS channel analysis.
This work addresses the lack of intuitive geometric interpretation in classical semiparametric efficiency theory, which has hindered the derivation and understanding of influence functions. The paper reformulates the theory within a differential geometric framework on the space of probability distributions, drawing an analogy to multivariate calculus: statistical paths, scores, and influence functions correspond respectively to curves, velocity vectors, and gradients. It demonstrates that the efficient influence function arises naturally as an orthogonal projection. By integrating functional analysis, differential geometry, and statistical inference, the study establishes a unified geometric interpretation of scores, tangent spaces, nuisance tangent spaces, and efficient influence functions. This synthesis not only clarifies several foundational theoretical issues but also substantially enhances the interpretability of methods in causal inference and missing data analysis.
This study addresses the long-standing misconception that estimator ranking inconsistencies in data attribution stem from approximation errors, revealing instead that they originate from counterfactual norm mismatches. By formalizing influence as a counterfactual estimator, this work establishes norm analysis as a necessary prerequisite for comparing estimators. It derives local decompositions to analytically characterize signal interaction mechanisms, validated through linearized approximations and controlled experiments. The primary contribution is the first demonstration that behavioral proxy selection critically impacts attribution quality, proving that differing norms directly induce ranking discrepancies. Furthermore, the proposed behavior-aligned norm successfully identifies target samples overlooked by default methods, substantially improving attribution accuracy.
This work addresses the computational and numerical challenges that commonly arise in practical implementations of higher-order influence function estimation, which often suffer from high-dimensional density estimation or inversion of large Gram matrices. The authors propose a stabilized estimation procedure that eliminates the need for sample splitting by incorporating Gram matrix regularization and a bilinear form structure. This approach avoids high-dimensional density estimation altogether while substantially improving numerical stability in finite samples. The method retains the theoretically optimal convergence rate and provides strong statistical guarantees alongside robust empirical performance, effectively overcoming key limitations of existing higher-order influence function estimators.
This work addresses the infeasibility of computing influence functions when model size vastly exceeds dataset size. The authors propose a dual representation–based approach to influence function estimation, incorporating kernel methods into influence analysis. Under the assumption of linearizable models, they explicitly construct a dual formulation that reduces computational complexity to scale with the dataset size rather than the number of model parameters. This framework represents the first efficient influence computation method whose complexity is governed by data size, enabling accurate estimation of the effect of removing individual data points on model parameters, predictions, and loss—particularly advantageous in large-model, small-data regimes where conventional approaches incur prohibitive computational costs.
Traditional influence functions fail in constrained learning settings because they neglect how data perturbations affect the feasible region, leading to biased estimates or infeasible solutions. This work proposes Directional Influence Functions (DIF), which explicitly incorporate constraints into the influence analysis framework for the first time. By modeling the optimality conditions of constrained optimization as a variational inequality and integrating sensitivity analysis with leave-one-out approximations, DIF accurately captures the effect of training sample perturbations on model parameters. Experiments on constrained linear regression and CNNs with fairness constraints demonstrate that DIF precisely replicates retraining results, significantly outperforming classical influence functions and their penalty-based variants, and exhibits strong alignment with actual retraining outcomes in predicting changes in test loss.