Score
Designs and analyzes statistical inference procedures that construct confidence intervals and confidence sets by leveraging optimal transport couplings or transport distances; builds algorithms to compute transport-based intervals and confidence regions and to adjust quantile-based intervals via transport mappings. Evaluates and optimizes coverage properties (including minimizing coverage deviation) and implements these methods for valid uncertainty quantification in finite-sample or otherwise challenging settings.
This work addresses the poor coverage performance of traditional confidence intervals in small-sample settings or with complex models, where reliance on asymptotic approximations often fails. The authors propose a novel confidence interval construction grounded in optimal transport theory, which minimizes coverage bias through optimal coupling and incorporates data-driven hyperparameter selection to enhance practical applicability. By moving beyond conventional quantile-based approaches, the method achieves substantially improved coverage accuracy and robustness across a range of estimation problems. The theoretical analysis rigorously establishes results concerning comparisons of probability measures, consistency, and finite-sample error bounds, providing a solid foundation for the proposed framework.
This work addresses the challenge of statistical inference under intractable likelihoods and composite contamination involving both geometric and total variation perturbations, where conventional methods often fail. The authors propose a robust optimal transport divergence grounded in empirical likelihood principles, enabling reliable parameter estimation and uncertainty quantification even when the model is misspecified but simulation-based data generation remains feasible. The approach integrates semi-discrete optimal transport, regularized bootstrap resampling, and stochastic subgradient optimization within a parallelizable simulation-based inference (SBI) framework that enjoys provable convergence guarantees. Theoretical analysis establishes the robustness of the proposed divergence under composite contamination, while experiments on challenging SBI benchmark tasks demonstrate its superior performance in terms of both robustness and statistical efficiency.
This paper addresses statistical inference for optimal transport (OT) maps by developing a nonparametric estimation framework and asymptotic theory grounded in empirical data. Methodologically, it integrates optimal transport theory, convergence analysis of probability measures, and large-sample statistical techniques to rigorously characterize estimation consistency and limiting distributions for diverse OT maps—including continuous, discrete, and regularized variants. The work introduces a unified inferential paradigm, providing the first rigorous asymptotic guarantees for point estimation, confidence band construction, and hypothesis testing of OT maps. These theoretical advances substantially enhance the interpretability and reliability of OT in practical applications such as causal inference, generative modeling, and distribution alignment. The resulting statistical toolkit—comprising estimators, uncertainty quantification procedures, and operational guidelines—is designed for cross-domain deployment and reproducible implementation. (149 words)
This study addresses the semi-discrete optimal transport problem under risk-sensitive settings by extending the classical objective of minimizing expected transport cost to minimizing a quantile of the transport cost distribution. By integrating quantile optimization, semi-discrete optimal transport theory, and geometric analysis, the work provides the first complete characterization of quantile-optimal transport plans that respect prescribed marginal constraints. An efficient simulation-based algorithm is proposed to compute such plans, and a novel tie-breaking rule is introduced to ensure solution uniqueness. Furthermore, the analysis reveals new geometric structures in spatial partitioning induced by the quantile objective, offering both theoretical foundations and computational tools for risk-sensitive transportation and regional segmentation.
This paper addresses the problem of finite-sample parameter inference for incompletely specified models. We propose the first confidence region construction method that guarantees exact finite-sample coverage probability equal to the nominal level. Our method employs a novel generalized Monte Carlo test statistic, built upon discrete optimal transport, to precisely characterize the sharp identification set. The resulting confidence region is computed via linear programming—ensuring computational feasibility and parameter independence while delivering exact coverage. To enhance efficiency, we design a conservative, consistent, and computationally efficient prescreening algorithm that substantially accelerates computation. Crucially, the method provides rigorous finite-sample validity: for any given sample size, the coverage probability is exactly equal to the pre-specified nominal level, without relying on asymptotic approximations or strong identification assumptions—thereby overcoming key limitations of conventional approaches.
Existing optimal transport (OT) mapping estimation theory heavily relies on Brenier’s theorem—which requires quadratic cost and absolutely continuous source distributions—rendering it inadequate for stochastic OT mappings with mass splitting, commonly encountered in real-world settings involving singular, discrete, or corrupted source/target distributions. Method: We propose a novel metric to quantify the quality of stochastic OT mappings and develop the first universal, robust, finite-sample optimal risk bound framework. Our approach integrates generalization error analysis, adversarially robust statistical learning, parameterized stochastic mapping modeling, and regularized empirical risk minimization. Contribution/Results: We derive near-optimal finite-sample risk bounds under minimal distributional assumptions. Experiments demonstrate substantial improvements in transport accuracy over conventional OT methods in challenging non-absolutely-continuous and corrupted-data regimes where standard approaches fail.
This study addresses the scalability and statistical validity bottlenecks in likelihood approximation and inference for complex simulation models by proposing a novel framework based on aggregated normalizing flow chains. Methodologically, it integrates information-theoretic formalization with sequential decision-making paradigms to construct flexible probability distributions through the sequential optimization of bijective transformation parameters. Furthermore, an empirical likelihood estimator under moment constraints is employed to iteratively update and aggregate the global flow parameters. This research establishes a surrogate model that simultaneously ensures computational feasibility and statistical power, enabling efficient parameter exploration, hypothesis testing, and uncertainty quantification. Ultimately, the proposed approach provides a reliable Bayesian inference solution for complex systems.
This work addresses the challenge of effectively integrating a small amount of coupled data with abundant uncoupled marginal observations to enhance downstream statistical inference. The authors propose a fully nonparametric approach that aligns marginal data with limited coupled samples via optimal transport projections and introduces an explicit estimator grounded in the notion of “shadow” couplings to extrapolate the dependence structure and improve estimation accuracy. The method offers geometric interpretability, numerical stability, and near-linear-time parallelizability. Theoretical guarantees are established by synthesizing tools from optimal transport theory, projection-based estimation, and sample complexity analysis. Extensive experiments on both synthetic and real-world datasets demonstrate the method’s high accuracy and computational efficiency.
This study addresses the computational bottleneck in evaluating the calibration of nested uncertainty sets within expensive simulation models. To overcome this limitation, the work proposes an efficient calibration assessment method grounded in a Bayesian framework and the Dirichlet-Multinomial model. By exploiting the nested structure, the approach directly processes interval outputs without requiring access to the full predictive distribution. Furthermore, it incorporates Bayes factor testing for statistical inference, substantially reducing the number of independent simulations needed. The proposed method successfully detects model miscalibration in data assimilation tasks under limited simulation budgets. Overall, this work significantly lowers computational costs while demonstrating both the effectiveness and practical utility of the proposed approach for calibrating complex simulation systems.
Traditional hybrid experimental designs struggle to robustly control the frequentist operating characteristics of Bayesian decisions under model misspecification and lack efficient sample size determination methods applicable to generalized posteriors. This work proposes a computationally efficient experimental design framework that requires simulations at only two sample sizes and leverages extrapolation modeling of posterior summary functions to infer performance across the entire sample size space. This approach enables identification of the minimal sample size and decision rule satisfying desired operating characteristics. It represents the first general and scalable method for sample size planning under generalized posteriors, substantially reducing computational burden while enhancing robustness to model misspecification. The method’s validity and broad applicability within Bayesian M-estimation–type experiments are demonstrated through the redesign of an adaptive clinical trial with time-to-event outcomes.