Score
Designs, implements, and evaluates models and estimators that map inputs (covariates or input distributions) to entire output probability distributions or to functions representing those distributions, including architectures that transform distributions into functional outputs. Builds procedures to fit and validate these distributional predictions, assess calibration and divergence, and support tasks such as counterfactual conditional-distribution construction and distributional comparisons over time.
Existing approaches often reduce functional responses to scalars or conditional mean curves, thereby failing to capture the full influence of covariates on the entire response distribution—including its shape, temporal dynamics, and variability. This work proposes a functional distributional random forest that uniquely integrates random forests with kernel methods in function spaces. By employing maximum mean discrepancy based on Sobolev kernels or operator-induced kernels at leaf nodes, the method estimates covariate-dependent full conditional distributions nonparametrically. It enables inference on arbitrary distributional functionals while preserving the realism of predicted samples. Simulations demonstrate the model’s ability to recover distributional dynamics missed by baseline methods, and an analysis of NHANES accelerometer data reveals significant covariate effects on both the median activity profiles and predictive dispersion.
This work proposes an additive nonlinear Bayesian regression framework to address the challenge of functional outputs in complex computer simulations that are jointly influenced by functional predictors defined over a fixed spatial domain and global scalar variables varying across simulation runs. The approach introduces a novel functional Gaussian process (fGP) prior that simultaneously models the unknown nonlinear effect of global variables across the entire spatial domain and captures spatially varying coefficients of local functional predictors through Gaussian processes. By explicitly encoding spatial dependence in the global effects, the fGP enables an interpretable decomposition of contributions from these two multiscale predictor types and provides principled uncertainty quantification. Experiments on synthetic data and the SLOSH hurricane storm surge model demonstrate that the method achieves high predictive accuracy alongside reliable uncertainty estimates.
This study addresses statistical inference for distributional models lacking classical density functions or finite moments by establishing a unified theoretical framework. The authors generalize Godambe’s inference functions to the space of distributions and introduce an observation operator to formally characterize diverse data-generating mechanisms, including point observations, interval censoring, and convolutional measurements. For the first time, inference functionals are defined in a distributional sense. Leveraging Schwartz’s theory of generalized functions, the Hájek–Le Cam convolution theorem, and Bhapkar–Godambe projections, the paper develops rigorous results on consistency, asymptotic normality, and optimality. A central contribution is the identification of a three-tier information hierarchy: Fisher information bounds the information captured by the observation operator, which in turn bounds the information attainable by any inference functional. The framework’s validity is demonstrated through applications to heavy-tailed distributions, interval-censored location models, and elliptical contour models.
This study addresses the challenge of jointly modeling calibration and control parameters in computer model calibration, where the distribution of calibration parameters is unknown while that of control parameters is known. To tackle this issue, the authors propose a nonparametric Bayesian calibration method based on measure decomposition. The approach preserves the known marginal distribution of the control parameters while employing stochastic process modeling and Bayesian inference to construct a posterior distribution over the input space that aligns with field observations. Notably, this work is the first within a nonparametric calibration framework to explicitly maintain the prior distributional properties of the control parameters, thereby substantially enhancing the physical consistency and scientific credibility of the calibration results.
When predictive models are deployed in new environments, their performance often degrades due to covariate shift and selective labeling, which jointly obscure accurate assessment of the true target risk. This work proposes an unbiased risk estimation method that integrates double machine learning with influence functions to simultaneously address both sources of bias for the first time. The approach is model-agnostic and compatible with general loss functions, constructing a corrected target risk estimator via nonparametric and plug-in techniques. Experiments on eICU electronic health record data demonstrate that the proposed method significantly outperforms baselines that handle only one type of bias or naively combine existing approaches, yielding more accurate tracking of the true target risk.
This study addresses the challenge of identifying and inferring integral functionals of conditional distributions with discontinuous outcomes by proposing a novel ReLU regression framework. The approach constructs closed-form estimators through covariate projection of ReLU-transformed outcome variables and recovers integrated conditional quantile functions via their convex conjugates obtained through the Legendre–Fenchel transform. By introducing ReLU activation into regression, this method enables direct identification of conditional distribution features under only weak distributional assumptions, accommodates discontinuous outcomes, and identifies average quantile treatment effects over arbitrary probability intervals. Leveraging Hadamard directional differentiability and the Delta method, the authors establish a unified theory for the consistent asymptotic distribution of the proposed estimators, substantially expanding the set of distributional parameters identifiable in empirical research.
This study addresses the unification of calibration concepts across classification and regression tasks, aiming to ensure consistency between predicted distributions and observed outcomes for diverse data types—continuous, discrete, nominal, and binary. The work introduces modal calibration for nominal outcomes and establishes a hierarchical framework distinguishing full, partial, and average calibration. It proposes a generalized definition of calibration based on predictive distribution functionals—such as means, quantiles, and event probabilities—and leverages probability integral transforms alongside constructive algorithms for analysis. Key contributions include demonstrating the logical independence between dual probability integral transform (PIT) calibration and existing discrete calibration notions, clarifying implication and independence relationships among various calibration types, and providing reproducible methods for generating illustrative examples and counterexamples.
This work addresses the challenge that existing proximal causal inference methods cannot identify the interventional joint distribution involving all proxy variables. To overcome this limitation, we propose, for the first time, an extended bridge function and establish a theoretical identifiability result for this distribution, integrating it into a kernel-based proximal causal inference framework. By introducing a synergistic mechanism between the interventional kernel and the extended bridge function, we derive novel identification conditions for the interventional joint distribution and develop a general, computable kernelized proximal identification algorithm. This contribution not only broadens the theoretical foundations of proximal causal inference but also provides a practical tool for estimating causal effects in settings with high-dimensional proxy variables.