transport map estimation

Design, build, and analyze estimators and algorithms that recover a transport map or a transport plan between probability distributions from finite data; establish their statistical properties (consistency, convergence rates, sample complexity) and fundamental limits (minimax lower bounds and hardness relations to optimal-transport formulations and related generative methods).

transportmapestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work investigates the statistical limits and sample complexity lower bounds for estimating any valid transport map without relying on optimal transport (OT). Formalizing the estimation task within a minimax framework, the study integrates tools from optimal transport theory, minimax analysis, and stability-based hypothesis testing to establish, for the first time, rigorous statistical lower bounds for non-optimal yet valid transport maps. The results demonstrate that under standard stability conditions, the statistical difficulty of estimating an arbitrary valid transport map is comparable to that of estimating the OT map itself. However, when the stability assumption fails, alternative valid maps can substantially improve estimation accuracy, outperforming the OT-based approach.

generative modelingminimax frameworkoptimal transport

Statistical Inference for Optimal Transport Maps: Recent Advances and Perspectives

Jun 23, 2025
SB
Sivaraman Balakrishnan
🏛️ Carnegie Mellon University | Massachusetts Institute of Technology

This paper addresses statistical inference for optimal transport (OT) maps by developing a nonparametric estimation framework and asymptotic theory grounded in empirical data. Methodologically, it integrates optimal transport theory, convergence analysis of probability measures, and large-sample statistical techniques to rigorously characterize estimation consistency and limiting distributions for diverse OT maps—including continuous, discrete, and regularized variants. The work introduces a unified inferential paradigm, providing the first rigorous asymptotic guarantees for point estimation, confidence band construction, and hypothesis testing of OT maps. These theoretical advances substantially enhance the interpretability and reliability of OT in practical applications such as causal inference, generative modeling, and distribution alignment. The resulting statistical toolkit—comprising estimators, uncertainty quantification procedures, and operational guidelines—is designed for cross-domain deployment and reproducible implementation. (149 words)

Developing limit theorems for transport map inferenceEstimating optimal transport maps from sample dataExtending results to special OT cases and variants

Estimation of Stochastic Optimal Transport Maps

Dec 10, 2025
SN
Sloan Nietert
🏛️ EPFL | Cornell University

Existing optimal transport (OT) mapping estimation theory heavily relies on Brenier’s theorem—which requires quadratic cost and absolutely continuous source distributions—rendering it inadequate for stochastic OT mappings with mass splitting, commonly encountered in real-world settings involving singular, discrete, or corrupted source/target distributions. Method: We propose a novel metric to quantify the quality of stochastic OT mappings and develop the first universal, robust, finite-sample optimal risk bound framework. Our approach integrates generalization error analysis, adversarially robust statistical learning, parameterized stochastic mapping modeling, and regularized empirical risk minimization. Contribution/Results: We derive near-optimal finite-sample risk bounds under minimal distributional assumptions. Experiments demonstrate substantial improvements in transport accuracy over conventional OT methods in challenging non-absolutely-continuous and corrupted-data regimes where standard approaches fail.

Develops a metric for evaluating stochastic optimal transport mapsExtends theory to real-world applications with stochastic transportProvides efficient estimators with robust finite-sample risk bounds

Minimax Rates of Estimation for Optimal Transport Map between Infinite-Dimensional Spaces

May 19, 2025
DP
Donlapark Ponnoprat
🏛️ Chiang Mai University | The University of Tokyo | RIKEN AIP

This paper addresses the nonparametric estimation of γ-smooth optimal transport maps in infinite-dimensional spaces, challenging the conventional belief that exponential sample complexity is necessary. We propose a constructive estimator combining regularized kernel smoothing with projection-based approximation, integrating tools from functional analysis, empirical process theory, and γ-smoothness characterization. Our analysis establishes, for the first time, the minimax optimal convergence rate of (O(n^{-1/(2+gamma)})), which is polynomial—significantly improving upon previously known exponential lower bounds. This yields the first computationally tractable and statistically optimal framework for estimating infinite-dimensional transport maps. Numerical experiments demonstrate that the proposed estimator substantially outperforms existing baselines on functional data tasks, achieving both statistical optimality and practical deployability.

Achieving polynomial minimax risk rates for γ-smooth mapsEstimating optimal transport maps in infinite-dimensional spacesOvercoming exponential data requirements for estimation

Score-based Generative Neural Networks for Large-Scale Optimal Transport

Oct 07, 2021
MD
Max Daniels
🏛️ Northeastern University | Brandeis University

To address the high computational cost and curse-of-dimensionality challenges in sampling optimal transport (OT) couplings for large-scale, high-dimensional data, this paper proposes an efficient learning framework based on score-based generative models. Specifically, conditioned on source samples, it iteratively generates target samples following the Sinkhorn-regularized OT coupling via Langevin dynamics. Crucially, it jointly parameterizes the score function and Sinkhorn potential functions—enabling, for the first time, end-to-end co-learning of score-based generation and OT coupling. We theoretically establish the convergence of gradient descent on the network parameters under mild assumptions. Experiments demonstrate that our method significantly improves both accuracy and speed of coupling estimation across diverse large-scale OT tasks, while maintaining scalability and practical applicability.

Learning Sinkhorn coupling via score-based networksSampling optimal transport coupling between distributionsSolving high-dimensional transport without linear programming

Latest Papers

What's happening recently
View more

The massive use of Machine Learning (ML) tools in industry comes with critical challenges, such as the lack of explainable models and the use of black-box algorithms. We address this issue by applying Optimal Transport theory in the analysis of responses of ML models to variations in the distribution of input variables. We find the closest distribution, in the Wasserstein sense, that satisfies a given constraintt and examine its impact on model behavior. Furthermore, we establish convergence results for this projected distribution and demonstrate our approach using examples and real-world datasets in both regression and classification settings.

This study investigates whether empirical subgradients of sample-based optimal transport objectives converge to the subdifferential of the population objective, thereby ensuring that subgradient methods consistently approximate population stationary points. By leveraging subdifferential analysis and graphical convergence theory, the work establishes—for the first time—the graphical convergence of empirical subgradients within the optimal transport framework. It further reveals the critical role of parametric smoothness in balancing statistical consistency and optimization stability, showing that nonsmooth settings may induce derivative instability even with large samples. The theoretical findings are validated in applications including risk-averse optimization, fairness-constrained learning, and sliced Wasserstein problems, demonstrating that standard subgradient methods indeed converge consistently to population stationary points.

nonsmooth optimizationoptimal transportpopulation objective

This work addresses the challenge of effectively integrating a small amount of coupled data with abundant uncoupled marginal observations to enhance downstream statistical inference. The authors propose a fully nonparametric approach that aligns marginal data with limited coupled samples via optimal transport projections and introduces an explicit estimator grounded in the notion of “shadow” couplings to extrapolate the dependence structure and improve estimation accuracy. The method offers geometric interpretability, numerical stability, and near-linear-time parallelizability. Theoretical guarantees are established by synthesizing tools from optimal transport theory, projection-based estimation, and sample complexity analysis. Extensive experiments on both synthetic and real-world datasets demonstrate the method’s high accuracy and computational efficiency.

coupled datadata integrationmarginal data

This work addresses the poor coverage performance of traditional confidence intervals in small-sample settings or with complex models, where reliance on asymptotic approximations often fails. The authors propose a novel confidence interval construction grounded in optimal transport theory, which minimizes coverage bias through optimal coupling and incorporates data-driven hyperparameter selection to enhance practical applicability. By moving beyond conventional quantile-based approaches, the method achieves substantially improved coverage accuracy and robustness across a range of estimation problems. The theoretical analysis rigorously establishes results concerning comparisons of probability measures, consistency, and finite-sample error bounds, providing a solid foundation for the proposed framework.

asymptotic approximationconfidence intervalscoverage probability

Hot Scholars

YM

Youssef Marzouk

Professor, Massachusetts Institute of Technology
computational mathematicsuncertainty quantificationinverse problemsdata assimilation
RB

Ricardo Baptista

University of Toronto
uncertainty quantificationinverse problemsdata assimilationcomputational statistics
TS

Taiji Suzuki

The University of Tokyo
StatisticsMachine learning
MK

Matthias Katzfuss

Professor of Statistics, University of Wisconsin–Madison
Spatio-Temporal StatisticsGaussian ProcessesUQProbabilistic Machine Learning
ZO

Zijing Ou

Imperial College London
machine learning